Claude Sonnet vs Claude Haiku — which Claude is enough for the job?
Same family, two different jobs. Claude Haiku is half the price ($1 / $5 vs $2 / $10 per 1M) and the fastest tier on Anthropic's own latency table. Claude Sonnet is clearly stronger on hard work — 85.2 vs 73.3 on SWE-bench Verified, with a 1M context against 200K. Choose Sonnet as the default for coding and multi-step reasoning; choose Haiku for high-volume classification, extraction, and routing.
| Category | Winner | Margin |
|---|---|---|
| Agentic coding · SWE-bench Verified | AClaude Sonnet | Sonnet 85.2 vs Haiku 73.3 — a 12-point gap (llm-stats, Jul 2026) |
| Overall intelligence · AA Intelligence Index | AClaude Sonnet | Sonnet 53 vs Haiku 30 on the 9-eval composite (Artificial Analysis v4.1) |
| Raw speed · output tokens/s | BClaude Haiku | Anthropic rates Haiku "Fastest" and Sonnet "Fast"; AA measures ~100 vs ~75 tok/s |
| API price · per 1M tokens | BClaude Haiku | $1 / $5 vs $2 / $10 — Haiku is exactly half on both sides while intro pricing lasts |
| Real cost on prose · tokenizer-adjusted | BClaude Haiku | Sonnet 5 runs Anthropic's newer tokenizer, so the gap per 1M characters is nearer 2.7× |
| Context window · API model | AClaude Sonnet | 1M vs 200K tokens — five times the room for repos and long documents |
| Max output · per response | AClaude Sonnet | 128K vs 64K tokens; Sonnet also gets the 300K batch-output beta |
| Knowledge freshness · reliable cutoff | AClaude Sonnet | Sonnet 5 is reliable to Jan 2026; Haiku 4.5 stops at Feb 2025 |
| Price stability · what happens Sep 1 | BClaude Haiku | Sonnet's $2 / $10 is introductory through Aug 31; Haiku has no announced change |
| Best overall | ·Depends | Sonnet is the default for real work, Haiku is the volume tier — most teams run both |
If the task needs reasoning, not just throughput.
- Coding — 85.2 vs 73.3 on SWE-bench Verified, a 12-point gap on real repository tasks
- Context — a 1M-token window against Haiku's 200K, five times the room for repos and long documents
- Longer answers — 128K max output per response versus 64K, plus a 300K batch-output beta Haiku doesn't get
- Fresher knowledge — reliable to Jan 2026 against Haiku's Feb 2025 cutoff
- Adaptive thinking — Sonnet 5 scales its own reasoning effort per request instead of a fixed thinking switch
If the work is simple and there is a lot of it.
- Half price — $1 / $5 per 1M against Sonnet's $2 / $10, and $0.50 / $2.50 on the Batch API
- Fastest tier — Anthropic's own latency table rates Haiku 4.5 the fastest Claude; Artificial Analysis measures ~100 output tokens/s versus ~75
- Cheaper prose — the older tokenizer means roughly 2.7× lower cost per 1M characters, not 2×
- Cheaper cache — $0.10/M cache reads and $1.25/M five-minute writes, half of Sonnet's
- Enough for the easy 80% — 73.3 on SWE-bench Verified covers boilerplate, extraction, classification, and routing
| Aspect | Claude Sonnet | Claude Haiku |
|---|---|---|
| API · inputper 1M tokens · from snapshot | $2.00 | $1.00 B wins |
| API · outputper 1M tokens · from snapshot | $10.0 | $5.00 B wins |
| Cached inputPrompt-cache read $/1M · from snapshot | $0.20 | $0.10 B wins |
| Effective API costBlended workload $/1M · from snapshot | $1.36 | $0.68 B wins |
| API context windowMax input tokens · from snapshot | 1M A wins | 200K |
| Real cost / 1M charsTokenizer-adjusted prose — Sonnet 5 runs the newer, token-hungrier tokenizer | $0.77 | $0.29 B wins |
| Standard rate from Sep 1, 2026What the API bill looks like after the intro period | $3 / $15 per 1M Sonnet 5's $2 / $10 is introductory pricing through Aug 31, 2026; Anthropic's pricing table already lists the September rates, batch included ($1.50 / $7.50) | $1 / $5 per 1M No repricing announced for Haiku 4.5 — so the same-family gap widens from 2× to 3× on September 1 B wins |
| Consumer plan accessWhere each model shows up in Claude subscriptions | The default model Sonnet is what the Claude apps hand you on Free, Pro ($17/mo billed annually, $20 monthly) and Max — you don't have to pick it A wins | Selectable, every tier claude.com/pricing lists Haiku under Models and usage on Free, Pro, Max and Team seats; nobody buys a subscription for it — it earns its keep on the API |
| Capability | Claude Sonnet | Claude Haiku |
|---|---|---|
| API context window | 1M tokens | 200K tokens |
| Max output per response | 128K tokens | 64K tokens |
| 300K batch-output beta | ✓ output-300k header | ✗ |
| Positioning | Balanced default | Volume / low-latency tier |
| Comparative latency (vendor table) | Fast | Fastest |
| Measured output speed | ~75 tokens/s | ~100 tokens/s |
| Agentic coding (SWE-bench Verified) | 85.2% | 73.3% |
| Intelligence Index (AA v4.1) | 53 | 30 |
| Adaptive thinking | ✓ | ✗ |
| Extended thinking toggle | ✗ (adaptive instead) | ✓ thinking.type |
| Vision / image input | ✓ | ✓ |
| Prompt-cache read | ✓ $0.20/M | ✓ $0.10/M |
| 5-minute cache write | $2.50/M | $1.25/M |
| Batch API (50% off) | ✓ $1 / $5 | ✓ $0.50 / $2.50 |
| Tokenizer generation | ~ Newer, ~35% more tokens | ✓ Previous, leaner |
| Reliable knowledge cutoff | Jan 2026 | Feb 2025 |
| Tool use / agents | ✓ | ✓ |
| MCP support | ✓ Native | ✓ Native |
| Cloud availability | API, Bedrock, Vertex, Foundry | API, Bedrock, Vertex, Foundry |
| Consumer app availability | ✓ Default model | ✓ Selectable |
The numbers, not the spin.
Claude Sonnet
The model the Claude apps hand you by default — strong enough for most production work, and cheap enough that few teams bother routing around it.
Strengths
- Coding — 85.2 on SWE-bench Verified, twelve points clear of Haiku on real repository tasks
- Reasoning — 53 on the Artificial Analysis Intelligence Index against Haiku's 30
- Context — a full 1M-token window at standard pricing, five times Haiku's 200K
- Output room — 128K tokens per response, and up to 300K through the Batch API beta
- Recency — reliable knowledge through Jan 2026, a year fresher than Haiku
Weaknesses
- Twice Haiku's token price today, and three times it once the introductory rate expires on Aug 31, 2026
- The newer Anthropic tokenizer produces roughly a third more tokens for the same text (2.6 vs 3.5 characters per token on our calibration), so the real gap on prose is wider than the sticker gap
- Slower per token — Artificial Analysis measures ~75 output tokens/s against Haiku's ~100
- Overkill for classification, extraction, and routing, where the extra capability changes nothing
Best for
- Coding agents working inside a real repository
- Multi-step reasoning and tool-heavy workflows
- Long documents, large diffs, and 200K+ token contexts
- Anything where a wrong answer costs more than the tokens
Claude Haiku
The volume tier — half the price, the fastest Claude, and good enough for the shallow, repetitive calls that make up most of an agent's traffic.
Strengths
- Price — $1 / $5 per 1M, exactly half Sonnet's current rate and a third of it from September
- Speed — Anthropic's own latency table puts Haiku 4.5 ahead of every other Claude
- Tokenizer — the previous, leaner tokenizer means fewer billed tokens for the same text
- Cache economics — $0.10/M reads and $1.25/M five-minute writes are half Sonnet's on both sides
- Batch floor — $0.50 / $2.50 per 1M is the cheapest first-party Claude you can buy
Weaknesses
- Twelve points behind Sonnet on SWE-bench Verified and 23 points behind on the intelligence index
- 200K context and 64K max output — a fifth and a half of Sonnet's ceilings
- Knowledge stops at Feb 2025, so it is the wrong model for anything recency-sensitive
- No adaptive thinking; only the older fixed extended-thinking toggle
- Excluded from the 300K batch-output beta
Best for
- Classification, extraction, tagging, and routing at volume
- Support-ticket triage and templated drafting
- Sub-agent and tool-call steps inside a larger Sonnet-led pipeline
- Batch jobs where latency and unit cost dominate
Classifying a million support emails a month
You need each message tagged by intent and urgency. The prompt is short, the output is a label, and the volume is relentless.
Reasoning: Nothing in this task rewards deeper reasoning — it rewards throughput and unit cost. Haiku is half the price per token, roughly 2.7× cheaper per 1M characters of real prose once the tokenizer difference is counted, and the fastest Claude on Anthropic's own latency table. Sonnet would produce the same labels for more money.
Coding agent working inside a real repository
The agent reads files, edits code, runs tests, and iterates until the suite is green. Failed attempts cost tokens and your attention.
Reasoning: The 12-point SWE-bench Verified gap (85.2 vs 73.3) lands exactly here: Haiku retries on tasks Sonnet closes, and a retry costs more than the price difference. Sonnet's 1M context also holds far more of the repository at once. Haiku is a false economy on repo work.
Support-ticket triage at 10,000 tickets a day
Each conversation is short and formulaic; you need a first-pass draft reply and a routing decision, with a human reviewing the edge cases.
Reasoning: Anthropic's own worked example puts a support conversation at roughly 3,700 tokens and prices 10,000 of them at about $37 on Haiku 4.5. The same volume on Sonnet costs twice that today and three times from September, for output a human reviewer will edit anyway. Run Haiku as the first pass and escalate whatever the confidence score flags.
Reading a 400-page contract in one pass
You want cross-references resolved across the whole document, not a chunked summary stitched together from a retrieval index.
Reasoning: Haiku's 200K window cannot hold the document, so you would be back to chunking and losing the cross-references you came for. Sonnet's 1M window fits it whole at standard pricing, and 128K max output leaves room for a real clause-by-clause write-up. Context, not intelligence, decides this one.
Cost-capped agent with a per-step router
You run a fixed monthly API budget and your agent makes thousands of calls a day, most of them shallow tool-calls and formatting steps.
Reasoning: The biggest lever on this bill is not the model — it is the split. Routing the shallow majority to Haiku at $1 / $5 and reserving Sonnet for the steps that actually reason typically halves spend without moving output quality, because the cheap steps were never capability-bound. Haiku carries the volume here by design.
Claude Pro subscriber deciding what to click
You pay $20/mo (or $17 billed annually) and the model picker offers Haiku alongside Sonnet and the Opus-class models.
Reasoning: On a subscription you are spending message allowance, not dollars per token, so Haiku's price advantage buys you nothing — and its Feb 2025 knowledge cutoff and 200K context make it the weaker choice for chat. Leave the default alone. Haiku is a model you buy on the API, not one you pick in the app.
Frequently asked.
Common questions about this comparison, with sources where they matter.
Q · 01 Is Claude Sonnet or Haiku better? +
Q · 02 Is Haiku really half the price of Sonnet? +
$1 / $5 vs $2 / $10 per 1M tokens. Two things change that. First, Sonnet 5's $2 / $10 is introductory pricing through Aug 31, 2026 — Anthropic's pricing table already lists $3 / $15 from September 1, which turns the gap into 3×. Second, Sonnet 5 uses Anthropic's newer tokenizer, which produces roughly a third more tokens for the same text (2.6 vs 3.5 characters per token on our calibration), so on ordinary prose the real gap is closer to 2.7× even now. See the real-cost-per-1M-characters row above, and model your own mix in the LLM API cost calculator.Q · 03 Is Haiku good enough for coding? +
Q · 04 How should I split traffic between them? +
Q · 05 Which one is faster? +
100 output tokens/s for Haiku against 75 for Sonnet at max effort. The time-to-first-token difference is larger still, because Sonnet's adaptive thinking spends time reasoning before it emits anything. For interactive UI latency, Haiku is the noticeably snappier model.Q · 06 Do they have the same context window? +
1M input tokens and 128K max output; Haiku 4.5 carries 200K input and 64K output. Sonnet is also the only one of the two eligible for the 300K batch-output beta. If your prompts routinely exceed 200K tokens, the choice is already made for you — and note that Sonnet's newer tokenizer means its 1M window holds fewer characters per token than the raw number suggests.