GPT-6 Astra vs Claude Fable 5.1 — identical stickers, opposite economics
Two frontier models launched two days apart at the same $10 and $50 per million. Fable 5.1 reads cache at a quarter of OpenAI's rate, so per token it is 9% cheaper — and it beats Astra on almost every eval. But running Artificial Analysis's whole index cost $8,523 on Fable 5.1 against $3,013 on Astra, because it thinks far longer. Same price, and a 2.8x difference in what the work costs.
| Category | Winner | Margin |
|---|---|---|
| List price · input and output per million | ·Identical | $10 in and $50 out on both cards, and both halve to $5 and $25 on batch — the sticker cannot separate them |
| Cache read rate · what re-sent context costs | BClaude Fable 5.1 | $0.25 against $1.00 — Anthropic prices this model's cache hit at 2.5% of base input, against the 10% every other Claude model and OpenAI both use |
| Effective rate · our 92/8 agentic blend, 82% cache | BClaude Fable 5.1 | $6.26 against $6.82 per million — a 9% edge that comes entirely from the cache-read line, since every other rate matches |
| Cost of finishing the work · running all of AA's index | AGPT-6 Astra | $3,013 against $8,523 — Fable 5.1 at max effort spends $6,239 on reasoning tokens against Astra's $1,824, and the cheaper rate cannot cover that |
| Overall intelligence · AA Intelligence Index v4.1.1 | BClaude Fable 5.1 | 65.7 against 61.2 — four and a half points of quality between two models whose rate cards are identical to the cent |
| Agentic work · AA Agentic Index | BClaude Fable 5.1 | 61.3 against 51.5 — nearly ten points, on the workload that decides most frontier-model purchases |
| Knowledge reliability · AA-Omniscience, hallucination-penalised | ·Dead heat | 43.45 against 43.40 — five hundredths of a point apart, and the only measure on which these two are genuinely equal |
| Science QA · GPQA Diamond | AGPT-6 Astra | 96.1% against 93.7% — one of only two evals in the index that Astra wins |
| Terminal work · Terminal-Bench v2.1 | BClaude Fable 5.1 | 91.4% against 88.4% — both are strong here, and both are near the top of the whole index |
| Availability · where you can call it | BClaude Fable 5.1 | Breadth, not gating: Fable 5.1 runs on the Claude API, Bedrock, Vertex, Microsoft Foundry and AWS. Astra opened to everyone on OpenAI's own API within three days of launch, but its Foundry route stays behind a Limited Access Program |
| Cache flexibility · how long a cache entry lives | BClaude Fable 5.1 | Anthropic sells a one-hour cache write at $20 alongside the five-minute one at $12.50; OpenAI publishes a single cache-write rate |
| Speed · output tokens per second | ·Unmeasured | Fable 5.1 runs 66.5 tokens a second at 7.29 minutes of decode per task; AA has published no speed figure for Astra yet |
| Best overall | ·Depends | Fable 5.1 if you are buying answers and can get an account today; Astra if you are buying tokens and your reasoning budget is the constraint |
If what you cannot afford is thinking time.
- 2.8x cheaper to finish the job — $3,013 against $8,523 to run Artificial Analysis's whole index, at the same list price
- A third of the reasoning bill — $1,824 against $6,239 of reasoning tokens on that run, which is the entire difference
- Level on knowledge reliability — 43.40 against 43.45 on AA-Omniscience, a dead heat with a model that spends three times as much getting there
- Best science QA of the two — 96.1% on GPQA Diamond against 93.7%
- Effort you control — five discrete levels from low to max, against adaptive thinking that is always on
- Two published rate cards — OpenAI's and Microsoft Foundry's, agreeing cell for cell
If you want the better model and you want it today.
- Wins the index — 65.7 against 61.2, and it leads on GDPval, τ³-Banking, Terminal-Bench v2.1, SciCode, AA-LCR and Humanity's Last Exam
- Ten points of agentic headroom — 61.3 against 51.5 on AA's Agentic Index
- A quarter of the going rate — $0.25 per million, 2.5% of base input, where Anthropic's own footnote puts every other Claude model at 10% and OpenAI charges that same 10% on Astra
- 9% cheaper per token on our blend — $6.26 against $6.82 on the 92/8 agentic mix at an 82% cache-hit rate
- Buyable now — Claude API, Amazon Bedrock, Google Vertex, Microsoft Foundry and the Claude Platform on AWS
- A one-hour cache tier — $20 per million to write, which OpenAI does not offer at all
| Aspect | GPT-6 Astra | Claude Fable 5.1 |
|---|---|---|
| API · inputper 1M tokens · from snapshot | $10.00 | $10.00 |
| API · outputper 1M tokens · from snapshot | $50.00 | $50.00 |
| API · cached inputper 1M cached tokens | $1.00 | $0.25 B wins |
| Effective costblended 92/8 · 82% cache | $6.82 | $6.26 B wins |
| Cache read as a share of inputthe only rate that differs | 10% of input $1 per million against a $10 base — the standard ratio across OpenAI's 5.x and 6 families | 2.5% of input $0.25 per million. Anthropic's pricing page carries a footnote saying cache reads cost 10% of base input on every Claude model except Fable 5.1 and Mythos 5.1, where it is 2.5% B wins |
| Measured cost of the workrunning all of AA's Intelligence Index | $3,013 At max effort: $909 of input, $2,105 of output, of which $1,824 is reasoning tokens A wins | $8,523 At max effort with default fallback — Anthropic's most expensive configuration: $1,483 of input, $7,041 of output, of which $6,239 is reasoning. A lower effort setting would cut this sharply, and so would the score |
| Cache write tierswhat it costs to fill the cache | $12.50 · 5 min One tier, billed at 1.25x the uncached input rate. There is no longer-lived option to buy | $12.50 · 5 min · $20 · 1 hr Two tiers. The one-hour write costs 60% more up front and is the cheaper choice for any agent whose context survives longer than five minutes between turns B wins |
| Long-context surchargevery large prompts | 2x above 272K Prompts over 272,000 input tokens reprice the entire request at 2x input and cache rates and 1.5x output — $20, $2 and $75 | No published step Anthropic publishes long-context and per-platform pricing separately rather than as a single threshold multiplier on this card B wins |
| Who can call itaccess route | One open route OpenAI's own API, open to any account since September 6 after three days of Trusted-Access-only launch. Microsoft's Foundry resale stays behind its own Limited Access Program, so the second route is still gated | Five platforms, today Claude API, Amazon Bedrock, Google Vertex AI, Microsoft Foundry and the Claude Platform on AWS, all with published model ids B wins |
| Capability | GPT-6 Astra | Claude Fable 5.1 |
|---|---|---|
| Released | Sep 3, 2026 | Sep 1, 2026 |
| Vendor's own positioning | "Our most capable model" — opens the models index | For demanding reasoning and long-horizon agentic work; Opus 5 is the default |
| Access | ✓ OpenAI API — Foundry route still gated | ✓ Generally available on five platforms |
| Input price | $10 / 1M | $10 / 1M |
| Output price | $50 / 1M | $50 / 1M |
| Cache read | $1.00 · 10% of input | $0.25 · 2.5% of input |
| Batch | $5 / $25 — half of standard | $5 / $25 — half of standard |
| One-hour cache tier | ✗ Not offered | ✓ $20 / 1M to write |
| Context window | 1,050,000 | 1,000,000 |
| Maximum output tokens | 128,000 | 128,000 |
| Knowledge cutoff | Apr 30, 2026 | June 2026 |
| Reasoning control | Five levels: low → max | Adaptive thinking, always on, default effort high |
| Retirement commitment | None published | ✓ No sooner than Sep 1, 2027 |
| AA Intelligence Index v4.1.1 | 61.2 | 65.7 |
| AA Agentic Index | 51.5 | 61.3 |
| AA-Omniscience | 43.40 | 43.45 |
| GDPval-AA v2, normalised | 56.5 | 67.7 |
| Terminal-Bench v2.1 | 88.4% | 91.4% |
| SciCode | 54.1% | 62.0% |
| τ³-Banking | 41.4% | 47.2% |
| Humanity's Last Exam | 54.7% | 59.1% |
| GPQA Diamond | 96.1% | 93.7% |
| CritPt | 31.7% | 29.7% |
| AA-LCR, long-context reasoning | 74.3% | 80.0% |
| Reasoning spend on AA's index | $1,824 | $6,239 |
| Output speed | Not yet published | 66.5 tokens/sec |
| Second vendor rate card | ✓ Microsoft Foundry, matching cell for cell | ✓ Bedrock, Vertex and Foundry all list it |
The numbers, not the spin.
GPT-6 Astra
The same sticker as Anthropic's flagship, reached with a third of the thinking.
Strengths
- 2.8x cheaper per finished job — $3,013 against $8,523 to run Artificial Analysis's whole index, on identical list rates
- Reasoning discipline — $1,824 of reasoning tokens across that index against Fable 5.1's $6,239
- Equal on knowledge reliability — 43.40 against 43.45 on AA-Omniscience, a dead heat reached far more cheaply
- Best science QA of the two — 96.1% on GPQA Diamond, and a narrow win on CritPt
- Explicit effort control — five levels from low to max, so you can buy less thinking when the task is easy
- Slightly larger window — 1,050,000 tokens against 1,000,000, with a 922,000-token input ceiling
Weaknesses
- Behind on the index — 61.2 against 65.7, losing GDPval, τ³-Banking, Terminal-Bench, SciCode, AA-LCR and HLE
- Ten points behind on agentic work — 51.5 against 61.3, at the same price per token
- Cache read costs 4x more — $1 against $0.25 per million, on cards that are otherwise identical
- One platform against five — Astra is OpenAI's API, and Microsoft's Foundry resale is still gated; Fable 5.1 ships on the Claude API, Bedrock, Vertex, Foundry and AWS with published ids on each
- A hard long-context cliff — above 272,000 input tokens the entire request reprices at 2x input and 1.5x output
- No one-hour cache — five-minute writes only, which suits chat better than a long-running agent
Best for
- Reasoning-heavy work where the token budget, not the score, is the constraint
- Science and technical Q&A
- Shops standardised on OpenAI tooling and the Responses API
Claude Fable 5.1
The better model at the same list price, and the more expensive one to actually run.
Strengths
- Wins the index by 4.5 points — 65.7 against 61.2, the top score on Artificial Analysis's board
- Agentic lead — 61.3 against 51.5, plus wins on GDPval, τ³-Banking, Terminal-Bench v2.1, SciCode, AA-LCR and Humanity's Last Exam
- Cache read at 2.5% of base input — $0.25 per million, against the 10% Anthropic's own pricing footnote applies to every other Claude model, and the 10% OpenAI charges on Astra
- 9% cheaper per token on our blend — $6.26 against $6.82 on the 92/8 agentic mix at an 82% cache-hit rate
- Buy it anywhere — Claude API, Bedrock, Vertex, Microsoft Foundry and the Claude Platform on AWS
- A retirement date you can plan against — Anthropic commits to no sooner than September 1, 2027
Weaknesses
- 2.8x the bill to finish the same suite — $8,523 against $3,013, driven by $6,239 of reasoning tokens
- Thinking is not optional — adaptive reasoning is always on with a default effort of high, so the cheap configuration has to be asked for
- Slower — 66.5 output tokens a second and 7.29 minutes of decode per index task
- Loses science QA — 93.7% on GPQA Diamond against 96.1%, and CritPt goes the same way
- The vendor's own headline is optimistic — Anthropic advertises up to 45% cheaper for highly agentic workloads against Fable 5, but on our standard blend the saving is 8.3%
Best for
- Long-horizon agents and multi-step coding work
- Anyone who needs the highest score available and can buy it today
- Workloads that re-send a large stable context on every turn
You want the best model available and can pay for it
Quality decides the purchase; the token bill is a rounding error against the salary of the person waiting for the output.
Reasoning: Fable 5.1. It leads the Intelligence Index 65.7 to 61.2, leads the Agentic Index 61.3 to 51.5, and wins six of the nine index components at the same list price. The two stickers are identical, so this is not a case of paying more for more — it is the same $10/$50 buying a lower score. What separates the bills is the cache read, $0.25 against $1, which pulls the same-sticker models about 9% apart on a cached workload.
A high-volume pipeline with a fixed monthly budget
Thousands of calls a day, a hard ceiling on spend, and quality that only has to clear a bar rather than top a leaderboard.
Reasoning: Astra, and it is not close on the numbers that matter here. The rates are identical, but the measured cost of finishing Artificial Analysis's index was $3,013 against $8,523 — a 2.8x difference produced entirely by how long each model thinks. At volume that is the whole budget. Set effort to low or medium and the gap widens further, because Astra exposes the control and Fable 5.1's thinking is always on.
An agent that re-sends a large stable context every turn
A long system prompt, a big repository map, dozens of turns an hour, and a cache that either holds or does not.
Reasoning: Fable 5.1, on the cache line alone. Its cache read is $0.25 per million against $1 — 2.5% of base input against 10% — and it sells a one-hour cache write at $20 that OpenAI does not offer at all. On our 92/8 blend at an 82% hit rate that is a 9% saving before you count the ten-point agentic lead. Astra's five-minute-only cache is the wrong shape for an agent that pauses between turns.
Answering hard science questions
Graduate-level physics, chemistry and biology, where the answer is checkable and being confidently wrong is the failure mode.
Reasoning: Astra, narrowly and for a specific reason. It takes GPQA Diamond 96.1% to 93.7% and CritPt 31.7% to 29.7%, and it matches Fable 5.1 on AA-Omniscience — the hallucination-penalised measure — at a third of the reasoning spend. Both models are near the ceiling on GPQA, so treat this as a tie-breaker rather than a verdict, and note that Fable 5.1 still wins Humanity's Last Exam.
Procurement will not onboard another vendor
The work has to run under a cloud contract you already hold — AWS, Google Cloud or Azure — rather than a direct account with the model vendor.
Reasoning: Fable 5.1, and not narrowly. It ships on Amazon Bedrock, Google Vertex AI, Microsoft Foundry and the Claude Platform on AWS with a published model id on each, so the bill arrives through paper that is already signed. Astra's only open route is a direct OpenAI account: its Foundry listing exists but sits behind a Limited Access Program, so holding an Azure contract does not by itself get you the model. Where the constraint is contractual rather than technical, it decides the question before any benchmark gets a say.
Very large single prompts
Whole codebases, discovery document sets, transcript archives — inputs that run past a quarter of a million tokens.
Reasoning: Watch the cliff before the score. Astra reprices the entire request at 2x input and 1.5x output above 272,000 input tokens, so a 300K prompt costs double rather than slightly more; Fable 5.1 publishes no equivalent single-threshold multiplier on its card. Fable 5.1 also leads AA-LCR, the long-context reasoning eval, 80.0% to 74.3%. Both arguments point the same way here.
Frequently asked.
Common questions about this comparison, with sources where they matter.