The 20x in Claude Code caching is real — it's a penalty, not a discount
A cache miss costs exactly 20x the warm read of the same history. What that does to your plan limits, which settings change it, and the number nobody can compute.
8 posts in this category.
Concrete how-to guides for working with LLM APIs and tooling -- enabling prompt caching the right way, instrumenting cost telemetry, picking a model for a specific workload class. Each guide includes runnable code or commands, the exact vendor page that documents the API surface, and a date stamp that lets the reader judge how fresh the recipe is. We do not publish a how-to until it has been run against a live API at least once. The category opens empty in Phase D. The first guides will follow once the analysis catalogue covers the conceptual ground a how-to builds on.
A cache miss costs exactly 20x the warm read of the same history. What that does to your plan limits, which settings change it, and the number nobody can compute.
Showing 7 of 8 posts — page 1 of 1
Batch halves token rates at OpenAI, Anthropic and Google -- but takes 20% at xAI and nothing on its flagship. When batch beats a warm cache, and when it costs more.
Prompt caching fails silently -- no error, no warning, just a full-price bill. The floors, the breakpoint rules, and everything that quietly invalidates a cache.
Five levers, one deployment order: which context feature to reach for by symptom, what each does to your cache and your bill, and one config that runs them all.
Five MCP servers burn ~55k tokens before you ask anything. Tool search and programmatic tool calling (both now GA) cut that 85%+ — with one caveat that bites.
Anthropic's compaction API summarizes an agent's history when it hits a token threshold. How it works, the billing pass you don't see, and when it backfires.
Anthropic's context editing clears stale tool results from an agent's window, cutting token use up to 84%. How it works, the config, and the prompt-cache catch.
Caveman mode strips filler from Claude's output. The 65% headline is real for output tokens. Here's what it actually saves on your monthly bill.