Last verified

GPT-6 Astra vs Claude Fable 5.1 — identical stickers, opposite economics

Two frontier models launched two days apart at the same $10 and $50 per million. Fable 5.1 reads cache at a quarter of OpenAI's rate, so per token it is 9% cheaper — and it beats Astra on almost every eval. But running Artificial Analysis's whole index cost $8,523 on Fable 5.1 against $3,013 on Astra, because it thinks far longer. Same price, and a 2.8x difference in what the work costs.

§ 01 / VERDICT

Who wins, category by category.

Skip to decision tree →
Category Winner Margin
List price · input and output per million ·Identical $10 in and $50 out on both cards, and both halve to $5 and $25 on batch — the sticker cannot separate them
Cache read rate · what re-sent context costs BClaude Fable 5.1 $0.25 against $1.00 — Anthropic prices this model's cache hit at 2.5% of base input, against the 10% every other Claude model and OpenAI both use
Effective rate · our 92/8 agentic blend, 82% cache BClaude Fable 5.1 $6.26 against $6.82 per million — a 9% edge that comes entirely from the cache-read line, since every other rate matches
Cost of finishing the work · running all of AA's index AGPT-6 Astra $3,013 against $8,523 — Fable 5.1 at max effort spends $6,239 on reasoning tokens against Astra's $1,824, and the cheaper rate cannot cover that
Overall intelligence · AA Intelligence Index v4.1.1 BClaude Fable 5.1 65.7 against 61.2 — four and a half points of quality between two models whose rate cards are identical to the cent
Agentic work · AA Agentic Index BClaude Fable 5.1 61.3 against 51.5 — nearly ten points, on the workload that decides most frontier-model purchases
Knowledge reliability · AA-Omniscience, hallucination-penalised ·Dead heat 43.45 against 43.40 — five hundredths of a point apart, and the only measure on which these two are genuinely equal
Science QA · GPQA Diamond AGPT-6 Astra 96.1% against 93.7% — one of only two evals in the index that Astra wins
Terminal work · Terminal-Bench v2.1 BClaude Fable 5.1 91.4% against 88.4% — both are strong here, and both are near the top of the whole index
Availability · where you can call it BClaude Fable 5.1 Breadth, not gating: Fable 5.1 runs on the Claude API, Bedrock, Vertex, Microsoft Foundry and AWS. Astra opened to everyone on OpenAI's own API within three days of launch, but its Foundry route stays behind a Limited Access Program
Cache flexibility · how long a cache entry lives BClaude Fable 5.1 Anthropic sells a one-hour cache write at $20 alongside the five-minute one at $12.50; OpenAI publishes a single cache-write rate
Speed · output tokens per second ·Unmeasured Fable 5.1 runs 66.5 tokens a second at 7.29 minutes of decode per task; AA has published no speed figure for Astra yet
Best overall ·Depends Fable 5.1 if you are buying answers and can get an account today; Astra if you are buying tokens and your reasoning budget is the constraint
CHOOSE A · GPT-6 ASTRA

If what you cannot afford is thinking time.

  • 2.8x cheaper to finish the job — $3,013 against $8,523 to run Artificial Analysis's whole index, at the same list price
  • A third of the reasoning bill — $1,824 against $6,239 of reasoning tokens on that run, which is the entire difference
  • Level on knowledge reliability — 43.40 against 43.45 on AA-Omniscience, a dead heat with a model that spends three times as much getting there
  • Best science QA of the two — 96.1% on GPQA Diamond against 93.7%
  • Effort you control — five discrete levels from low to max, against adaptive thinking that is always on
  • Two published rate cards — OpenAI's and Microsoft Foundry's, agreeing cell for cell
CHOOSE B · CLAUDE FABLE 5.1

If you want the better model and you want it today.

  • Wins the index — 65.7 against 61.2, and it leads on GDPval, τ³-Banking, Terminal-Bench v2.1, SciCode, AA-LCR and Humanity's Last Exam
  • Ten points of agentic headroom — 61.3 against 51.5 on AA's Agentic Index
  • A quarter of the going rate — $0.25 per million, 2.5% of base input, where Anthropic's own footnote puts every other Claude model at 10% and OpenAI charges that same 10% on Astra
  • 9% cheaper per token on our blend — $6.26 against $6.82 on the 92/8 agentic mix at an 82% cache-hit rate
  • Buyable now — Claude API, Amazon Bedrock, Google Vertex, Microsoft Foundry and the Claude Platform on AWS
  • A one-hour cache tier — $20 per million to write, which OpenAI does not offer at all
§ 02 / PRICING

What it actually costs.

Cost calculator →
Aspect GPT-6 Astra Claude Fable 5.1
API · inputper 1M tokens · from snapshot even verified Sep 02 $10.00 $10.00
API · outputper 1M tokens · from snapshot even verified Sep 02 $50.00 $50.00
API · cached inputper 1M cached tokens verified Sep 02 $1.00 $0.25 B wins
Effective costblended 92/8 · 82% cache verified Sep 02 $6.82 $6.26 B wins
Cache read as a share of inputthe only rate that differs verified Sep 04 10% of input $1 per million against a $10 base — the standard ratio across OpenAI's 5.x and 6 families 2.5% of input $0.25 per million. Anthropic's pricing page carries a footnote saying cache reads cost 10% of base input on every Claude model except Fable 5.1 and Mythos 5.1, where it is 2.5% B wins
Measured cost of the workrunning all of AA's Intelligence Index verified Sep 04 $3,013 At max effort: $909 of input, $2,105 of output, of which $1,824 is reasoning tokens A wins $8,523 At max effort with default fallback — Anthropic's most expensive configuration: $1,483 of input, $7,041 of output, of which $6,239 is reasoning. A lower effort setting would cut this sharply, and so would the score
Cache write tierswhat it costs to fill the cache verified Sep 04 $12.50 · 5 min One tier, billed at 1.25x the uncached input rate. There is no longer-lived option to buy $12.50 · 5 min · $20 · 1 hr Two tiers. The one-hour write costs 60% more up front and is the cheaper choice for any agent whose context survives longer than five minutes between turns B wins
Long-context surchargevery large prompts verified Sep 04 2x above 272K Prompts over 272,000 input tokens reprice the entire request at 2x input and cache rates and 1.5x output — $20, $2 and $75 No published step Anthropic publishes long-context and per-platform pricing separately rather than as a single threshold multiplier on this card B wins
Who can call itaccess route verified Sep 04 One open route OpenAI's own API, open to any account since September 6 after three days of Trusted-Access-only launch. Microsoft's Foundry resale stays behind its own Limited Access Program, so the second route is still gated Five platforms, today Claude API, Amazon Bedrock, Google Vertex AI, Microsoft Foundry and the Claude Platform on AWS, all with published model ids B wins
§ 03 / FEATURES

Feature-by-feature, side by side.

Download CSV →
Capability GPT-6 Astra Claude Fable 5.1
Released Sep 3, 2026 Sep 1, 2026
Vendor's own positioning "Our most capable model" — opens the models index For demanding reasoning and long-horizon agentic work; Opus 5 is the default
Access ✓ OpenAI API — Foundry route still gated ✓ Generally available on five platforms
Input price $10 / 1M $10 / 1M
Output price $50 / 1M $50 / 1M
Cache read $1.00 · 10% of input $0.25 · 2.5% of input
Batch $5 / $25 — half of standard $5 / $25 — half of standard
One-hour cache tier ✗ Not offered ✓ $20 / 1M to write
Context window 1,050,000 1,000,000
Maximum output tokens 128,000 128,000
Knowledge cutoff Apr 30, 2026 June 2026
Reasoning control Five levels: low → max Adaptive thinking, always on, default effort high
Retirement commitment None published ✓ No sooner than Sep 1, 2027
AA Intelligence Index v4.1.1 61.2 65.7
AA Agentic Index 51.5 61.3
AA-Omniscience 43.40 43.45
GDPval-AA v2, normalised 56.5 67.7
Terminal-Bench v2.1 88.4% 91.4%
SciCode 54.1% 62.0%
τ³-Banking 41.4% 47.2%
Humanity's Last Exam 54.7% 59.1%
GPQA Diamond 96.1% 93.7%
CritPt 31.7% 29.7%
AA-LCR, long-context reasoning 74.3% 80.0%
Reasoning spend on AA's index $1,824 $6,239
Output speed Not yet published 66.5 tokens/sec
Second vendor rate card ✓ Microsoft Foundry, matching cell for cell ✓ Bedrock, Vertex and Foundry all list it
§ 04 / BENCHMARKS

The numbers, not the spin.

Overall intelligence · AA Intelligence Index v4.1.1
GPT-6 Astra
61.2%
Claude Fable 5.1
65.7%
Artificial Analysis Intelligence Index v4.1.1 — nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR · Astra at max effort, Fable 5.1 at max effort with default fallback · read Sep 4, 2026
Agentic work · AA Agentic Index
GPT-6 Astra
51.5%
Claude Fable 5.1
61.3%
Same source, same configurations, same date · nearly ten points on the workload that decides most frontier-model purchases, at an identical list price
Knowledge reliability · AA-Omniscience Index
GPT-6 Astra
43.4%
Claude Fable 5.1
43.5%
Same source and date · rewards correct answers, penalises hallucinations, no penalty for refusing · five hundredths of a point apart. Note what it costs each of them to get there: Astra spends $1,824 of reasoning tokens across the index, Fable 5.1 spends $6,239
Science QA · GPQA Diamond
GPT-6 Astra
96.1%
Claude Fable 5.1
93.7%
Index component, same source and date · one of only two components Astra wins, and both models sit near the ceiling — treat it as a tie-breaker, not a verdict
Terminal work · Terminal-Bench v2.1
GPT-6 Astra
88.4%
Claude Fable 5.1
91.4%
Index component, same source and date · both rank near the top of the whole index here, which is why the agentic gap between them is easier to see in AA's composite Agentic Index than in any single terminal benchmark
Hard exams · Humanity's Last Exam
GPT-6 Astra
54.7%
Claude Fable 5.1
59.1%
Index component, same source and date · the hardest closed-form set in the index. Fable 5.1's lead here is consistent with its lead on the composite, and it comes with the reasoning bill described above
§ 05 / DEEP DIVE

What each does best.

Brand hubs →
A · OPENAI

GPT-6 Astra

The same sticker as Anthropic's flagship, reached with a third of the thinking.

Strengths

  • 2.8x cheaper per finished job — $3,013 against $8,523 to run Artificial Analysis's whole index, on identical list rates
  • Reasoning discipline — $1,824 of reasoning tokens across that index against Fable 5.1's $6,239
  • Equal on knowledge reliability — 43.40 against 43.45 on AA-Omniscience, a dead heat reached far more cheaply
  • Best science QA of the two — 96.1% on GPQA Diamond, and a narrow win on CritPt
  • Explicit effort control — five levels from low to max, so you can buy less thinking when the task is easy
  • Slightly larger window — 1,050,000 tokens against 1,000,000, with a 922,000-token input ceiling

Weaknesses

  • Behind on the index — 61.2 against 65.7, losing GDPval, τ³-Banking, Terminal-Bench, SciCode, AA-LCR and HLE
  • Ten points behind on agentic work — 51.5 against 61.3, at the same price per token
  • Cache read costs 4x more — $1 against $0.25 per million, on cards that are otherwise identical
  • One platform against five — Astra is OpenAI's API, and Microsoft's Foundry resale is still gated; Fable 5.1 ships on the Claude API, Bedrock, Vertex, Foundry and AWS with published ids on each
  • A hard long-context cliff — above 272,000 input tokens the entire request reprices at 2x input and 1.5x output
  • No one-hour cache — five-minute writes only, which suits chat better than a long-running agent

Best for

  • Reasoning-heavy work where the token budget, not the score, is the constraint
  • Science and technical Q&A
  • Shops standardised on OpenAI tooling and the Responses API
B · ANTHROPIC

Claude Fable 5.1

The better model at the same list price, and the more expensive one to actually run.

Strengths

  • Wins the index by 4.5 points — 65.7 against 61.2, the top score on Artificial Analysis's board
  • Agentic lead — 61.3 against 51.5, plus wins on GDPval, τ³-Banking, Terminal-Bench v2.1, SciCode, AA-LCR and Humanity's Last Exam
  • Cache read at 2.5% of base input — $0.25 per million, against the 10% Anthropic's own pricing footnote applies to every other Claude model, and the 10% OpenAI charges on Astra
  • 9% cheaper per token on our blend — $6.26 against $6.82 on the 92/8 agentic mix at an 82% cache-hit rate
  • Buy it anywhere — Claude API, Bedrock, Vertex, Microsoft Foundry and the Claude Platform on AWS
  • A retirement date you can plan against — Anthropic commits to no sooner than September 1, 2027

Weaknesses

  • 2.8x the bill to finish the same suite — $8,523 against $3,013, driven by $6,239 of reasoning tokens
  • Thinking is not optional — adaptive reasoning is always on with a default effort of high, so the cheap configuration has to be asked for
  • Slower — 66.5 output tokens a second and 7.29 minutes of decode per index task
  • Loses science QA — 93.7% on GPQA Diamond against 96.1%, and CritPt goes the same way
  • The vendor's own headline is optimistic — Anthropic advertises up to 45% cheaper for highly agentic workloads against Fable 5, but on our standard blend the saving is 8.3%

Best for

  • Long-horizon agents and multi-step coding work
  • Anyone who needs the highest score available and can buy it today
  • Workloads that re-send a large stable context on every turn
§ 06 / SCENARIOS

Picked by scenario.

More scenarios →
01

You want the best model available and can pay for it

Quality decides the purchase; the token bill is a rounding error against the salary of the person waiting for the output.

Reasoning: Fable 5.1. It leads the Intelligence Index 65.7 to 61.2, leads the Agentic Index 61.3 to 51.5, and wins six of the nine index components at the same list price. The two stickers are identical, so this is not a case of paying more for more — it is the same $10/$50 buying a lower score. What separates the bills is the cache read, $0.25 against $1, which pulls the same-sticker models about 9% apart on a cached workload.

Picked
Claude Fable 5.1
Runner-up: Astra, if your own evals favour it — its knowledge-reliability profile differs from the composite
02

A high-volume pipeline with a fixed monthly budget

Thousands of calls a day, a hard ceiling on spend, and quality that only has to clear a bar rather than top a leaderboard.

Reasoning: Astra, and it is not close on the numbers that matter here. The rates are identical, but the measured cost of finishing Artificial Analysis's index was $3,013 against $8,523 — a 2.8x difference produced entirely by how long each model thinks. At volume that is the whole budget. Set effort to low or medium and the gap widens further, because Astra exposes the control and Fable 5.1's thinking is always on.

Picked
GPT-6 Astra
Runner-up: Fable 5.1 at a reduced effort setting, if you can measure the quality you lose
03

An agent that re-sends a large stable context every turn

A long system prompt, a big repository map, dozens of turns an hour, and a cache that either holds or does not.

Reasoning: Fable 5.1, on the cache line alone. Its cache read is $0.25 per million against $1 — 2.5% of base input against 10% — and it sells a one-hour cache write at $20 that OpenAI does not offer at all. On our 92/8 blend at an 82% hit rate that is a 9% saving before you count the ten-point agentic lead. Astra's five-minute-only cache is the wrong shape for an agent that pauses between turns.

Picked
Claude Fable 5.1
Runner-up: Astra, if your turns are close enough together that a five-minute cache never lapses
04

Answering hard science questions

Graduate-level physics, chemistry and biology, where the answer is checkable and being confidently wrong is the failure mode.

Reasoning: Astra, narrowly and for a specific reason. It takes GPQA Diamond 96.1% to 93.7% and CritPt 31.7% to 29.7%, and it matches Fable 5.1 on AA-Omniscience — the hallucination-penalised measure — at a third of the reasoning spend. Both models are near the ceiling on GPQA, so treat this as a tie-breaker rather than a verdict, and note that Fable 5.1 still wins Humanity's Last Exam.

Picked
GPT-6 Astra
Runner-up: Fable 5.1 for anything that mixes science with long-horizon work
05

Procurement will not onboard another vendor

The work has to run under a cloud contract you already hold — AWS, Google Cloud or Azure — rather than a direct account with the model vendor.

Reasoning: Fable 5.1, and not narrowly. It ships on Amazon Bedrock, Google Vertex AI, Microsoft Foundry and the Claude Platform on AWS with a published model id on each, so the bill arrives through paper that is already signed. Astra's only open route is a direct OpenAI account: its Foundry listing exists but sits behind a Limited Access Program, so holding an Azure contract does not by itself get you the model. Where the constraint is contractual rather than technical, it decides the question before any benchmark gets a say.

Picked
Claude Fable 5.1
Runner-up: Astra direct, if procurement can be moved for a single vendor
06

Very large single prompts

Whole codebases, discovery document sets, transcript archives — inputs that run past a quarter of a million tokens.

Reasoning: Watch the cliff before the score. Astra reprices the entire request at 2x input and 1.5x output above 272,000 input tokens, so a 300K prompt costs double rather than slightly more; Fable 5.1 publishes no equivalent single-threshold multiplier on its card. Fable 5.1 also leads AA-LCR, the long-context reasoning eval, 80.0% to 74.3%. Both arguments point the same way here.

Picked
Claude Fable 5.1
Runner-up: Astra under 272K tokens, where its larger 1,050,000 window is otherwise unused anyway

Frequently asked.

Common questions about this comparison, with sources where they matter.

Q · 01 They cost the same. Which one is actually cheaper? +
It depends on what you buy. Per token, Fable 5.1, by about 9%: the two cards are identical at $10 input and $50 output, but Anthropic reads cache at $0.25 against OpenAI's $1, which pulls our blended figure to $6.26 against $6.82. Per finished task, Astra, by a lot: Artificial Analysis spent $3,013 running its whole index on Astra and $8,523 on Fable 5.1, because Fable 5.1 at max effort burned $6,239 of reasoning tokens against Astra's $1,824. A rate card ranks models by price; a workload ranks them by how much thinking they do.
Q · 02 Why is Anthropic's cache read so much cheaper? +
Because Anthropic made an explicit exception for this generation. Its pricing page footnote says prompt cache reads cost 10% of the base input price — "2.5% on Claude Fable 5.1 and Claude Mythos 5.1". Base input, cache writes and output are all unchanged from the Fable 5 card at $10, $12.50 and $50; the cache read is the only rate that moved, from $1 to $0.25. Anthropic advertises up to 45% cheaper for highly agentic workloads on the strength of it. On this site's standard 92/8 blend at an 82% cache-hit rate the saving against Fable 5 works out at 8.3% — the vendor's figure needs a far higher read-to-write ratio and a more input-heavy mix than our blend assumes. Both numbers are correct; the assumption is the difference.
Q · 03 Which is the better model? +
Fable 5.1, on the neutral board and not narrowly. It scores 65.7 against 61.2 on Artificial Analysis's Intelligence Index v4.1.1 and 61.3 against 51.5 on the Agentic Index, and it wins GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, AA-LCR and Humanity's Last Exam. Astra wins two of the nine components — GPQA Diamond (96.1% against 93.7%) and CritPt — and ties AA-Omniscience at 43.40 against 43.45. Configuration matters when reading this: Astra ran at max effort, Fable 5.1 at max effort with default fallback, which is its most expensive setting.
Q · 04 Can I buy GPT-6 Astra right now? +
Yes, on OpenAI's own API — and that changed on September 6. Astra launched on September 3 to Trusted Access Program enterprises with general access promised in the coming days, and its model page now publishes rate limits from usage tier 1 upward, which an approval-gated model would not have. What has not changed is the breadth gap. Astra is one open route, and Microsoft still resells it through the Foundry Limited Access Program, so holding an Azure contract does not get you the model. Fable 5.1 runs on the Claude API, Amazon Bedrock, Google Vertex AI, Microsoft Foundry and the Claude Platform on AWS, with published ids on each. If your constraint is which cloud contract the bill lands on, that still settles it.
Q · 05 Does the context window matter here? +
Less than the pricing rule attached to it. Astra carries 1,050,000 tokens against Fable 5.1's 1,000,000, and both cap output at 128,000 — a 5% difference nobody will notice. What you will notice is that Astra reprices the entire request at 2x input and cache rates and 1.5x output once a prompt passes 272,000 input tokens, so a 300K-token call costs double rather than fractionally more. Anthropic publishes no equivalent single-threshold multiplier on this card. Fable 5.1 also leads AA-LCR, the long-context reasoning eval, 80.0% to 74.3%.
Q · 06 How can Astra be equal on hallucinations but behind everywhere else? +
AA-Omniscience measures something narrower than intelligence: it rewards correct answers, penalises confident wrong ones, and applies no penalty for refusing. A model can be excellent at knowing the boundary of its own knowledge while being worse at planning, tool use or multi-step work — which is exactly the shape of these two results. Astra scores 43.40 and Fable 5.1 43.45, a dead heat, while Fable 5.1 leads the Agentic Index by nearly ten points. If your product fails when the model invents a fact, that tie is the number to look at. If it fails when the model loses the thread across twenty tool calls, it is not.
Q · 07 Is Anthropic's cheaper cache enough to close the cost gap? +
No, and the arithmetic is worth seeing. The cache advantage is worth about $0.56 per million blended tokens — the whole of the 9% edge, since every other rate matches. The reasoning gap on Artificial Analysis's index was $4,415 in absolute dollars on a single suite. Caching discounts the input side; reasoning tokens are billed as output, at $50 per million on both cards. Any workload where the model thinks a lot will be dominated by the output line, and no cache rate reaches it. This is the same trap described in list price versus agent cost, running in the opposite direction.
Q · 08 What about GPT-5.6 Sol — is it still in the running? +
For agentic work, yes. Sol scores 57.8 on AA's Agentic Index — behind Fable 5.1's 61.3 but well ahead of Astra's 51.5 — at $4 and $20 per million, which is 2.5x cheaper than either model here. Its weakness is knowledge reliability, where it scores 22.0 on AA-Omniscience against roughly 43 for both models on this page. If your work is agent-shaped and your budget is real, the honest three-way answer is Sol; if it is answer-shaped and correctness matters more than throughput, it is one of these two. We take that pair apart in Astra against Sol.