Claude Opus 4.8 API Pricing
Claude Opus 4.8 is the Opus generation before Opus 5, still available at $5/M input, $25/M output and $0.50/M cache-hit input. Its successor ships on the identical rate card, so the migration question is behaviour, not price: Opus 5 thinks by default and bills those tokens as output. Pulled directly from platform.claude.com daily.
Run the numbers.
Live calculator pre-loaded with Claude Opus 4.8 rates. Tweak spend, output mix, or cache hit rate; share the URL to share the calculation.
Real-world presets.
Codebase-scale migration
Reading 100-page contracts
Support agent ticket triage
Research planning turn
Paste text. See tokens. See cost.
This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.
| Model | Input /M | Output /M | Effective blended | Context | Best for |
|---|---|---|---|---|---|
| Claude Opus 5 | $5.00 cache $0.50 | $25.00 | $3.41 same list price | 1M | The model that replaces 4.8 |
| Claude Opus 4.8 Current | $5.00 cache $0.50 | $25.00 | $3.41 agentic 92/8 | 1M | Previous Opus - same rate card |
| Claude Opus 4.7 | $5.00 cache $0.50 | $25.00 | $3.41 same list price | 1M | Previous Opus frontier tier |
| Claude Opus 4.6 | $5.00 cache $0.50 | $25.00 | $3.41 same list price | 1M | Older 1M Opus workloads |
| Claude Sonnet 4.6 | $3.00 cache $0.30 | $15.00 | $2.05 cheaper Anthropic sibling | 1M | Production agents and coding |
| Claude Haiku 4.5 | $1.00 cache $0.10 | $5.00 | $0.682 cheaper Anthropic sibling | 200K | Support and classification |
| GPT-5.5 | $5.00 cache $0.50 | $30.00 | $3.61 frontier competitor | 1M | OpenAI frontier agents |
| Gemini 3.1 Pro Preview | $2.00 cache $0.20 | $12.00 | $1.44 frontier preview competitor | 1M | Multimodal long-context analysis |
| DeepSeek V4 Pro | $1.32 cache $0.044 | $3.96 | $0.569 budget reasoning competitor | 1M | Low-cost reasoning workloads |
Frequently asked.
Practical Claude Opus 4.8 pricing questions, with regular usage separated from fast mode, cache writes, and batch discounts.
Q · 01 What is Claude Opus 4.8 priced at? +
$5/M input, $25/M output, and $0.50/M cache-hit input. Under AI//COST's 92/8 agentic blend with 82% cache hits, the effective planning figure is $3.41/M.Q · 02 Should I move to Claude Opus 5? +
$5/M input and $25/M output, the same cache and batch rates, and the same 1M context at standard pricing. The catch is billing behaviour - Opus 5 runs thinking by default, and those reasoning tokens bill at the output rate, so an identical workload can cost more on identical prices. Measure a sample before migrating high-volume traffic.Q · 03 How does prompt caching work for Opus 4.8? +
$0.50/M, which is 10% of the $5/M base input rate. Cache writes cost $6.25/M for 5-minute writes and $10/M for 1-hour writes. Those multipliers stack with data-residency modifiers.Q · 04 Does the full 1M context cost extra? +
1M token context window at standard pricing. A 900k-token request is billed at the same per-token rate as a 9k-token request.Q · 05 Is there a Batch API discount? +
$2.50/M input and $12.50/M output.Q · 06 What is fast mode pricing? +
$10/M input and $50/M output. Fast mode is not available with Batch API and is not available on Claude Platform on AWS.Q · 07 Does regional pricing differ? +
inference_geo: "us" applies a 1.1x multiplier to all token pricing categories. Global routing is the default and uses the standard prices shown here.Q · 08 How accurate is the tokenizer estimate? +
anthropic-bpe-estimate at 2.6 English characters per token. Anthropic says Opus 4.7 and later use a newer tokenizer that may use up to 35% more tokens for the same fixed text compared with previous models.