Running many agents at once: the tax is linear, the rebate is quadratic
Every extra agent pays for its own cold prefix. But re-reading history grows with the square of turns, so splitting one job across three agents ran 17% cheaper.
14 posts tagged "anthropic".
Every post tagged "anthropic" in the journal. Tag archives are auto-generated from post frontmatter -- one entry per unique tag across non-draft posts. Currently 14 posts share this tag. Use the category filter above to scope by editorial type.
Every extra agent pays for its own cold prefix. But re-reading history grows with the square of turns, so splitting one job across three agents ran 17% cheaper.
Showing 11 of 14 posts — page 1 of 2
A cache miss costs exactly 20x the warm read of the same history. What that does to your plan limits, which settings change it, and the number nobody can compute.
Batch halves token rates at OpenAI, Anthropic and Google -- but takes 20% at xAI and nothing on its flagship. When batch beats a warm cache, and when it costs more.
Prompt caching fails silently -- no error, no warning, just a full-price bill. The floors, the breakpoint rules, and everything that quietly invalidates a cache.
Ten routers launched on HN in five weeks. The peer-reviewed 2026 benchmarks say they barely beat a baseline. The arithmetic that decides it, on verified rates.
Five levers, one deployment order: which context feature to reach for by symptom, what each does to your cache and your bill, and one config that runs them all.
Cutting 63% of an agent's tokens moved its bill 9%. Cutting 1.2% moved it 15%. The arithmetic that converts token reductions into dollars, on verified 2026 rates.
Five MCP servers burn ~55k tokens before you ask anything. Tool search and programmatic tool calling (both now GA) cut that 85%+ — with one caveat that bites.
Anthropic's compaction API summarizes an agent's history when it hits a token threshold. How it works, the billing pass you don't see, and when it backfires.
Anthropic's context editing clears stale tool results from an agent's window, cutting token use up to 84%. How it works, the config, and the prompt-cache catch.
The $/M input/output sticker hides cache-write premiums, a tokenizer tax, context surcharges, invisible reasoning tokens, and more. Ten verified costs, with proofs.
In 2026 the biggest lever on your LLM bill isn't which model. It's how hard the model thinks: reasoning tokens bill as output, and most models think by default.