Last verified
NEWEST OPUS1M CONTEXTTEXT + VISIONTHINKING ON BY DEFAULTFAST MODE

Claude Opus 5 API Pricing

Claude Opus 5 launched July 24, 2026 on the same rate card as Opus 4.8: $5/M input, $25/M output, $0.50/M cache-hit input. The rates did not move, but the billing behaviour did — thinking is on by default, and reasoning tokens bill as output. Same list price, potentially larger invoice. Pulled directly from platform.claude.com daily.

Input - per 1M tokens
$5.00/M
Stable Opus tier same as 4.8
Output - per 1M tokens
$25.00/M
Includes thinking tokens same as 4.8
Cached input - 90% off
$0.50/M
Cache write $6.25 / $10 -90%
Effective - agentic blend
$3.41/M
92/8 split - 82% cache
§ 01 / TERMINAL

Run the numbers.

Live calculator pre-loaded with Claude Opus 5 rates. Thinking is on by default, so put your expected reasoning tokens in the reasoning field — they bill at the output rate and are the single biggest reason an Opus 5 invoice differs from an Opus 4.8 one at identical prices.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Estimate · anthropic-bpe-estimate · ≈2.6 chars/token Auto-counts as you type

This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.

Characters 666
Words 109
Tokens (estimated) 256 tokens
Cost as input · uncached $0.00128 USD
Cost as output · uncached $0.0064 USD
Cost as cached input $0.00013 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Claude Opus 5 Current $5.00 cache $0.50 $25.00 $3.41 agentic 92/8 1M Newest Opus - agentic coding and enterprise work
Claude Opus 4.8 $5.00 cache $0.50 $25.00 $3.41 same list price 1M The model Opus 5 replaces
Claude Fable 5 $10.00 cache $1.00 $50.00 $6.82 exactly 2x 1M Anthropic's top capability tier
Claude Sonnet 5 $2.00 cache $0.20 $10.00 $1.36 intro price to Aug 31 1M Speed-and-intelligence default
Claude Sonnet 4.6 $3.00 cache $0.30 $15.00 $2.05 cheaper sibling 1M Production agents on the old tokenizer
Claude Haiku 4.5 $1.00 cache $0.10 $5.00 $0.68 cheapest Claude 200K Support and classification volume
GPT-5.6 Sol $5.00 cache $0.50 $30.00 $3.81 frontier competitor 1.05M OpenAI frontier agents
Gemini 3.6 Flash $1.50 cache $0.15 $7.50 $0.96 cheaper competitor 1M Mainstream multimodal work
DeepSeek V4 Pro $0.43 cache $0.0036 $0.87 $0.14 budget reasoning 1M Low-cost reasoning workloads

Frequently asked.

Practical Claude Opus 5 pricing questions - what the rate card says, and where the real bill diverges from it.

Q · 01 What is Claude Opus 5 priced at? +
Anthropic's pricing page lists Claude Opus 5 at $5/M input, $25/M output and $0.50/M cache-hit input. Under AI//COST's 92/8 agentic blend with 82% cache hits, the effective planning figure is $3.41/M - identical to Opus 4.8.
Q · 02 Is Opus 5 more expensive than Opus 4.8? +
Not per token - every published rate is the same, including cache writes at $6.25/M and $10/M and batch at $2.50/M and $12.50/M. Your bill can still rise, because thinking is on by default on Opus 5 while Opus 4.8 ran without it unless asked. Reasoning tokens are billed as output at $25/M, so a request that produced 500 output tokens on 4.8 can produce several thousand billed tokens on 5. Model the reasoning tokens before migrating a high-volume workload.
Q · 03 How do I control the thinking cost? +
With the effort parameter, which Anthropic documents as the control for thinking depth on Opus 5. It supports low, medium, high, xhigh and max, and defaults to high on the Claude API. Anthropic says low and medium deliver strong quality at a fraction of the tokens. Thinking can also be disabled outright, but only at effort high or below - disabling it at xhigh or max returns a 400 error, which is a breaking change from Opus 4.8.
Q · 04 How does prompt caching work on Opus 5? +
Cache hits bill at $0.50/M, which is 10% of the base input rate. Writes bill at $6.25/M for the 5-minute TTL and $10/M for the 1-hour TTL, and they replace the base input price rather than adding to it. The minimum cacheable prompt is 512 tokens on Opus 5, down from 1,024 on Opus 4.8 - short system prompts that could never be cached before now can be, with no code change.
Q · 05 Does the full 1M context cost extra? +
No. Anthropic's pricing docs say the full 1M token window is included at standard pricing - a 900k-token request bills at the same per-token rate as a 9k-token one. On Opus 5 the 1M window is both the default and the maximum; there is no smaller-context variant to fall back to.
Q · 06 What do batch and fast mode cost? +
The Batch API halves both sides: $2.50/M input and $12.50/M output. Fast mode is a research preview at $10/M input and $50/M output - double the standard rate - and it is available on the Claude API only, not on Amazon Bedrock, Google Cloud or Microsoft Foundry. Fast mode cannot be combined with the Batch API.
Q · 07 How does it compare with Claude Fable 5 on price? +
Opus 5 is exactly half of Fable 5 on both sides of the rate card: $5 against $10 input, $25 against $50 output, and $3.41/M against $6.82/M on our blend. Anthropic positions Opus 5 as near-Fable-5 capability at half the cost; the price half is verifiable from the rate card, the capability half is a vendor claim we have not independently measured.
Q · 08 Does regional pricing differ? +
Yes. Setting inference_geo: "us" applies a 1.1x multiplier to every token category - input, output, cache writes and cache reads. Global routing is the default and uses the prices shown here. Amazon Bedrock and Google Cloud publish their own regional rates, with a 10% premium over global endpoints.
Q · 09 How accurate is the tokenizer estimate? +
The browser widget uses an anthropic-bpe-estimate at 2.6 English characters per token. Anthropic publishes no downloadable tokenizer, so this is an estimate. Opus 5 uses the tokenizer introduced with Opus 4.7, which produces roughly 30% more tokens for the same text than Sonnet 4.6 and earlier models - budget for that when comparing per-token prices across generations.