The list price is not the price your agent pays
The model 5.7% cheaper on the price tag billed an agent session 4.4x higher. The cached-read rate decides an agent bill, and it runs from 2% to 25% of input.
7 posts in this category.
Where the prose pulls apart pricing tables, agentic workload math, and cache economics that a single number on a vendor page cannot capture. Analysis posts here are anchored to verifiable data -- pricing snapshots from the canonical vendor page on the date of writing, worked examples whose arithmetic is auditable, and conclusions that move when the data moves. We publish these slowly. A single cornerstone analysis covers more ground than five thinkpieces, and we would rather correct one piece than retract three. Each entry carries its own verification date and the URLs of the sources it relied on, so the next person trying to repeat the math has a paper trail to follow.
The model 5.7% cheaper on the price tag billed an agent session 4.4x higher. The cached-read rate decides an agent bill, and it runs from 2% to 25% of input.
Showing 6 of 7 posts — page 1 of 1
Every extra agent pays for its own cold prefix. But re-reading history grows with the square of turns, so splitting one job across three agents ran 17% cheaper.
Ten routers launched on HN in five weeks. The peer-reviewed 2026 benchmarks say they barely beat a baseline. The arithmetic that decides it, on verified rates.
Cutting 63% of an agent's tokens moved its bill 9%. Cutting 1.2% moved it 15%. The arithmetic that converts token reductions into dollars, on verified 2026 rates.
The $/M input/output sticker hides cache-write premiums, a tokenizer tax, context surcharges, invisible reasoning tokens, and more. Ten verified costs, with proofs.
In 2026 the biggest lever on your LLM bill isn't which model. It's how hard the model thinks: reasoning tokens bill as output, and most models think by default.
Anthropic charges 10% of input for cache reads, with a 1.25x write fee. OpenAI auto-caches above 1024 tokens. The math changes which LLM is cheapest -- here is when.