The Batch API discount isn't 50% -- and it isn't always a discount
Batch halves token rates at OpenAI, Anthropic and Google -- but takes 20% at xAI and nothing on its flagship. When batch beats a warm cache, and when it costs more.
7 posts tagged "api-pricing".
Every post tagged "api pricing" in the journal. Tag archives are auto-generated from post frontmatter -- one entry per unique tag across non-draft posts. Currently 7 posts share this tag. Use the category filter above to scope by editorial type.
Batch halves token rates at OpenAI, Anthropic and Google -- but takes 20% at xAI and nothing on its flagship. When batch beats a warm cache, and when it costs more.
Showing 6 of 7 posts — page 1 of 1
Prompt caching fails silently -- no error, no warning, just a full-price bill. The floors, the breakpoint rules, and everything that quietly invalidates a cache.
Ten routers launched on HN in five weeks. The peer-reviewed 2026 benchmarks say they barely beat a baseline. The arithmetic that decides it, on verified rates.
Cutting 63% of an agent's tokens moved its bill 9%. Cutting 1.2% moved it 15%. The arithmetic that converts token reductions into dollars, on verified 2026 rates.
The $/M input/output sticker hides cache-write premiums, a tokenizer tax, context surcharges, invisible reasoning tokens, and more. Ten verified costs, with proofs.
In 2026 the biggest lever on your LLM bill isn't which model. It's how hard the model thinks: reasoning tokens bill as output, and most models think by default.
There is no single cheapest LLM API. The winner depends on your workload shape. Verified $/M prices for classification, chat, coding, RAG, batch, and reasoning.