The list price is not the price your agent pays
The model 5.7% cheaper on the price tag billed an agent session 4.4x higher. The cached-read rate decides an agent bill, and it runs from 2% to 25% of input.
18 posts tagged "cost-optimization".
Every post tagged "cost optimization" in the journal. Tag archives are auto-generated from post frontmatter -- one entry per unique tag across non-draft posts. Currently 18 posts share this tag. Use the category filter above to scope by editorial type.
The model 5.7% cheaper on the price tag billed an agent session 4.4x higher. The cached-read rate decides an agent bill, and it runs from 2% to 25% of input.
Showing 11 of 18 posts — page 1 of 2
Every extra agent pays for its own cold prefix. But re-reading history grows with the square of turns, so splitting one job across three agents ran 17% cheaper.
A cache miss costs exactly 20x the warm read of the same history. What that does to your plan limits, which settings change it, and the number nobody can compute.
Batch halves token rates at OpenAI, Anthropic and Google -- but takes 20% at xAI and nothing on its flagship. When batch beats a warm cache, and when it costs more.
Prompt caching fails silently -- no error, no warning, just a full-price bill. The floors, the breakpoint rules, and everything that quietly invalidates a cache.
Ten routers launched on HN in five weeks. The peer-reviewed 2026 benchmarks say they barely beat a baseline. The arithmetic that decides it, on verified rates.
Five levers, one deployment order: which context feature to reach for by symptom, what each does to your cache and your bill, and one config that runs them all.
Cutting 63% of an agent's tokens moved its bill 9%. Cutting 1.2% moved it 15%. The arithmetic that converts token reductions into dollars, on verified 2026 rates.
Five MCP servers burn ~55k tokens before you ask anything. Tool search and programmatic tool calling (both now GA) cut that 85%+ — with one caveat that bites.
Anthropic's compaction API summarizes an agent's history when it hits a token threshold. How it works, the billing pass you don't see, and when it backfires.
Anthropic's context editing clears stale tool results from an agent's window, cutting token use up to 84%. How it works, the config, and the prompt-cache catch.
Every legit way to get LLM compute for $0 in 2026: free API tiers, trial credits, student offers, and startup grants up to $350k — verified, with the catch on each.