The list price is not the price your agent pays
The model 5.7% cheaper on the price tag billed an agent session 4.4x higher. The cached-read rate decides an agent bill, and it runs from 2% to 25% of input.
18 posts across analysis, opinion, how-to, research, and roundups.
Independent analysis, opinion, and how-to coverage of LLM tooling -- written for engineers who pay the bill. No paid placements, no vendor talking points, no benchmarks lifted from press kits. Every model price comes from the vendor's own page, verified the week of publication; every cost calculation shows its math. The blog sits next to the price desk and the comparisons -- prose covers the questions a table cannot answer (when does a 90% cache discount stop mattering? which workload class breaks GPT vs Claude pricing parity?). Cornerstone analysis ships first; opinion, tools how-to, research summaries, and roundups will join as the catalogue grows. Read for the math, stay for the disagreements.
The model 5.7% cheaper on the price tag billed an agent session 4.4x higher. The cached-read rate decides an agent bill, and it runs from 2% to 25% of input.
Showing 11 of 18 posts — page 1 of 2
Every extra agent pays for its own cold prefix. But re-reading history grows with the square of turns, so splitting one job across three agents ran 17% cheaper.
A cache miss costs exactly 20x the warm read of the same history. What that does to your plan limits, which settings change it, and the number nobody can compute.
Batch halves token rates at OpenAI, Anthropic and Google -- but takes 20% at xAI and nothing on its flagship. When batch beats a warm cache, and when it costs more.
Prompt caching fails silently -- no error, no warning, just a full-price bill. The floors, the breakpoint rules, and everything that quietly invalidates a cache.
Ten routers launched on HN in five weeks. The peer-reviewed 2026 benchmarks say they barely beat a baseline. The arithmetic that decides it, on verified rates.
Five levers, one deployment order: which context feature to reach for by symptom, what each does to your cache and your bill, and one config that runs them all.
Cutting 63% of an agent's tokens moved its bill 9%. Cutting 1.2% moved it 15%. The arithmetic that converts token reductions into dollars, on verified 2026 rates.
Five MCP servers burn ~55k tokens before you ask anything. Tool search and programmatic tool calling (both now GA) cut that 85%+ — with one caveat that bites.
Anthropic's compaction API summarizes an agent's history when it hits a token threshold. How it works, the billing pass you don't see, and when it backfires.
Anthropic's context editing clears stale tool results from an agent's window, cutting token use up to 84%. How it works, the config, and the prompt-cache catch.
Every legit way to get LLM compute for $0 in 2026: free API tiers, trial credits, student offers, and startup grants up to $350k — verified, with the catch on each.