Last verified

GPT-6 Astra vs GPT-5.6 Sol — the new flagship against the one it was supposed to replace

OpenAI scaled every cell by the same factor: Astra costs exactly 2.5x Sol on input, cached input, cache writes, output and batch. The composite barely moves — 61.2 against 60.9 on Artificial Analysis's Intelligence Index — but the components split hard. Astra doubles Sol on knowledge reliability and loses to it on agentic work. Both are generally available since September 6, so the choice is now purely price against reasoning depth.

§ 01 / VERDICT

Who wins, category by category.

Skip to decision tree →
Category Winner Margin
Price per token · every rate, not just the headline BGPT-5.6 Sol Sol is 2.5x cheaper on input, cached input, cache writes, output and batch alike — the ratio is identical in all five cells
Overall intelligence · AA Intelligence Index v4.1.1 AGPT-6 Astra 61.2 against 60.9 — three tenths of a point for a 150% price increase, and inside the range where two runs can swap places
Agentic work · AA Agentic Index BGPT-5.6 Sol 57.8 against 51.5 — the older, cheaper model is 12% better at the workload most people buy a frontier model for
Knowledge reliability · AA-Omniscience, hallucination-penalised AGPT-6 Astra 43.4 against 22.0 — Astra's single decisive win, and the clearest thing the extra money buys
Hard exams · Humanity's Last Exam AGPT-6 Astra 54.7% against 49.5% — a real gap on the hardest closed-form questions in the index
Real-world professional tasks · GDPval-AA v2 BGPT-5.6 Sol 60.5 against 56.5 normalised — Sol is better at the eval built from actual occupational deliverables
Cost to finish the work · running AA's whole index BGPT-5.6 Sol $2,017 against $3,013 — Astra is only 1.49x the bill despite 2.5x the rates, because it is 20% less verbose
Verbosity · output tokens per index task AGPT-6 Astra 4,911 against 6,113 — Astra reaches the answer with a fifth fewer tokens, which is where the price premium partly disappears
Speed · decode time per index task BGPT-5.6 Sol Sol runs 76.5 output tokens a second and finishes a task in 3.73 minutes of decode; Astra is not yet published on either measure
Availability · can you actually call it today ·Depends Both generally available. Astra launched to Trusted Access enterprises only and opened within three days, so the availability gap that existed on launch week is gone
Price stability · how long the rate lasts AGPT-6 Astra Astra's card is standard; Sol's $4/$20 is promotional at least through Nov 21, 2026, and the pre-cut rate was $5/$30
Best overall ·Depends Astra for questions where a wrong confident answer is expensive; Sol for agents, for budgets, and for code you have to ship this week
CHOOSE A · GPT-6 ASTRA

If a confident wrong answer costs more than the tokens.

  • Knowledge reliability — 43.4 against 22.0 on AA-Omniscience, which rewards correct answers, penalises hallucinations and does not punish refusals
  • Hard exams — 54.7% on Humanity's Last Exam against 49.5%, and 96.1% on GPQA Diamond against 94.1%
  • Concise for its class — 4,911 output tokens per index task against Sol's 6,113, so the 2.5x rate turns into a 1.49x bill
  • Longer memory of the world — an April 30, 2026 knowledge cutoff against Sol's February 16
  • Vision headroom — 86.9% on MMMU-Pro against 83.4%
  • A price that is not a promotion — Astra's rate card is standard, while Sol's is explicitly temporary
CHOOSE B · GPT-5.6 SOL

If you are building agents, or paying the bill yourself.

  • Better at agentic work — 57.8 against 51.5 on AA's Agentic Index, and 60.5 against 56.5 on GDPval-AA v2
  • Exactly 2.5x cheaper — $4 and $20 against $10 and $50, with the same ratio on cached input, cache writes and batch
  • Cheaper on our blend too — $2.73 per million against $6.82 on the site's 92/8 agentic mix at an 82% cache-hit rate
  • Faster — 76.5 output tokens a second and 3.73 minutes of decode per index task
  • Its card is committed, not just current — OpenAI's footnote holds the $4/$20 promotional rate at least through November 21, 2026, while Astra's $10/$50 carries no such commitment
  • Nearly the same composite — 60.9 against 61.2, a gap smaller than most re-runs of the same benchmark
§ 02 / PRICING

What it actually costs.

Cost calculator →
Aspect GPT-6 Astra GPT-5.6 Sol
API · inputper 1M tokens · from snapshot verified Aug 23 $10.00 $4.00 B wins
API · outputper 1M tokens · from snapshot verified Aug 23 $50.00 $20.00 B wins
API · cached inputper 1M cached tokens verified Aug 23 $1.00 $0.40 B wins
Effective costblended 92/8 · 82% cache verified Aug 23 $6.82 $2.73 B wins
Measured cost of the workrunning all of AA's Intelligence Index verified Sep 04 $3,013 Artificial Analysis's own spend to run the full index on Astra at max effort: $909 of input against $2,105 of output, of which $1,824 is reasoning tokens $2,017 The same suite on Sol at max effort. Astra costs 2.5x per token but only 1.49x per suite, because it needs a fifth fewer output tokens to answer B wins
Is the price permanentpromotional or standard verified Sep 04 Standard card OpenAI publishes no promotional footnote for Astra. The rate is the rate, subject to the usual right to change it A wins Promotional to Nov 21 OpenAI's docs say Sol's promotional pricing runs at least through November 21, 2026. Before the cut the model listed at $5 and $30, so the gap to Astra would narrow to 2x input and 1.67x output if it reverts
Long-context surchargeprompts above 272K tokens verified Sep 04 2x in · 1.5x out Above 272,000 input tokens the whole request reprices: $20 input, $2 cached input, $25 cache writes, $75 output 2x in · 1.5x out Same rule, same threshold, applied to Sol's smaller card: $8 input, $0.80 cached input and $30 output on the full request
Discount routesbatch, flex and fast verified Sep 04 $5 / $25 batch Batch and Flex are 50% of standard; Fast mode is 2x, and Fast is unavailable for Astra under EU data residency. US data-zone endpoints carry a 1.1x uplift $2 / $10 batch The same 50% halving on the cheaper card, so the 2.5x ratio survives into batch as well. The 1.1x regional uplift applies here too B wins
Who can call itaccess route on day two verified Sep 06 Generally available Opened within three days of a Trusted-Access-only launch. The model page now publishes rate limits from usage tier 1 upward, which an approval-gated model would not have. Microsoft's Foundry route stays behind its own Limited Access Program Generally available Any API account, and the gpt-5.6 alias routes to it. Was the only callable option of the two during launch week
§ 03 / FEATURES

Feature-by-feature, side by side.

Download CSV →
Capability GPT-6 Astra GPT-5.6 Sol
Position in OpenAI's own docs Flagship — the models index now opens with it Held that line until Sep 3, 2026
Released Sep 3, 2026 Jul 9, 2026
Access ✓ Generally available — gated for its first three days ✓ Generally available
Second published rate card ✓ Microsoft Foundry, agreeing cell for cell OpenAI's own card only
Context window 1,050,000 1,050,000
Maximum input tokens 922,000 922,000
Maximum output tokens 128,000 128,000
Knowledge cutoff Apr 30, 2026 Feb 16, 2026
Reasoning effort levels low · medium · high · xhigh · max none · low · medium · high · xhigh · max
Input modalities Text, image Text, image
Chat Completions and Responses ✓ Both ✓ Both
Batch API ✓ At 50% of standard ✓ At 50% of standard
Realtime ✗ Not supported ✗ Not supported
Fine-tuning ✗ Not supported ✗ Not supported
Fast mode ~ 2x rates, and unavailable under EU data residency ~ 2x rates
Alias routing ✗ The Daybreak aliases still point at 5.6 ✓ The gpt-5.6 alias routes here
AA Intelligence Index v4.1.1 61.2 60.9
AA Agentic Index 51.5 57.8
AA-Omniscience 43.4 22.0
Terminal-Bench v2.1 88.4% 88.0%
GPQA Diamond 96.1% 94.1%
Humanity's Last Exam 54.7% 49.5%
SciCode 54.1% 56.1%
τ³-Banking 41.4% 44.3%
AA-LCR, long-context reasoning 74.3% 77.7%
MMMU-Pro 86.9% 83.4%
Output tokens per index task 4,911 6,113
Output speed Not yet published 76.5 tokens/sec
US data-zone uplift 1.1x 1.1x
§ 04 / BENCHMARKS

The numbers, not the spin.

Overall intelligence · AA Intelligence Index v4.1.1
GPT-6 Astra
61.2%
GPT-5.6 Sol
60.9%
Artificial Analysis Intelligence Index v4.1.1 — nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR · both at max effort · read Sep 4, 2026. Three tenths of a point separates a model that costs two and a half times as much
Agentic work · AA Agentic Index
GPT-6 Astra
51.5%
GPT-5.6 Sol
57.8%
Same source and date, max effort on both · the reversal that matters: OpenAI's new flagship is measurably worse than the model it displaced at the top of the docs, on the workload most buyers of a frontier model are actually running
Knowledge reliability · AA-Omniscience Index
GPT-6 Astra
43.4%
GPT-5.6 Sol
22.0%
Same source and date · rewards correct answers, penalises hallucinations, no penalty for refusing · Astra answers 62.6% correctly with a 51.3% hallucination rate on what it does attempt. This is the one axis where the price premium is unambiguous
Hard exams · Humanity's Last Exam
GPT-6 Astra
54.7%
GPT-5.6 Sol
49.5%
Index component, same source and date · HLE is the hardest closed-form set in the index, and it moves with the same thing AA-Omniscience measures — how much the model actually knows rather than how well it plans
Science QA · GPQA Diamond
GPT-6 Astra
96.1%
GPT-5.6 Sol
94.1%
Index component, same source and date · both are near the ceiling here, so read the two-point gap as confirmation of the knowledge story rather than as a decisive result
Professional tasks · GDPval-AA v2
GPT-6 Astra
56.5%
GPT-5.6 Sol
60.5%
Index component, normalised, same source and date · GDPval scores deliverables drawn from real occupations, and the cheaper model wins it — the second of the two components where paying 2.5x buys a worse result
§ 05 / DEEP DIVE

What each does best.

Brand hubs →
A · OPENAI

GPT-6 Astra

A frontier model priced as a step change and benchmarked as a sidestep.

Strengths

  • Knows more, and admits it less often — 43.4 against 22.0 on AA-Omniscience, the widest gap between these two on any measure
  • Hard-exam reasoning — 54.7% on Humanity's Last Exam and 96.1% on GPQA Diamond, both ahead of Sol
  • Concise — 4,911 output tokens per index task against 6,113, which is why 2.5x the rate is only 1.49x the bill
  • Fresher — knowledge cutoff April 30, 2026 against February 16
  • Two published rate cards — Microsoft's Foundry table agrees with OpenAI's cell for cell, and prints the regional uplift in dollars rather than as a multiplier
  • Effort ceiling — reasoning effort runs to max, and the ARC Prize write-up shows it building its own symbolic world models in unfamiliar environments

Weaknesses

  • Loses on agentic work — 51.5 against 57.8, and 56.5 against 60.5 on GDPval, both to the model it replaced
  • The composite does not pay for the premium — 61.2 against 60.9 on the Intelligence Index is inside the noise of a re-run, so the 2.5x buys a different component profile rather than a higher overall score
  • 2.5x on every rate — input, cached input, cache writes, output and batch all scale by the same factor, so no workload mix escapes the premium
  • Cache read is 10% of input — $1 against Anthropic's $0.25 on a same-priced model, which matters for anything that re-sends context
  • Fast mode is missing in the EU — OpenAI says to use Standard processing for EU data-residency requests
  • Its headline benchmark depends on the harness — ARC Prize measured 62.7% under its own scaffold and 99.9% under OpenAI's adapter, on the same set

Best for

  • Research and analysis where a confident wrong answer is expensive
  • Hard technical Q&A and exam-shaped reasoning
  • Long-horizon work that has to survive past 272K tokens of context
B · OPENAI

GPT-5.6 Sol

The model that held the flagship line for two months, and still wins the agentic half.

Strengths

  • Better agent — 57.8 on AA's Agentic Index against 51.5, plus wins on GDPval, τ³-Banking, SciCode and AA-LCR
  • 2.5x cheaper on every rate — $4 and $20 standard, $0.40 cached, $2 and $10 batch
  • $2.73 per million on our blend — against $6.82, on the 92/8 agentic mix at an 82% cache-hit rate
  • Fast — 76.5 output tokens a second, 3.73 minutes of decode per index task
  • Buyable — generally available, and the gpt-5.6 alias points at it
  • A composite you can round to equal — 60.9 against 61.2 on the same index

Weaknesses

  • Weak on knowledge reliability — 22.0 on AA-Omniscience against 43.4, the one place the gap is not close
  • Older world — a February 16, 2026 knowledge cutoff
  • Its price is a promotion — OpenAI says the rate holds at least through November 21, 2026, and the pre-cut card was $5 and $30
  • More verbose — 6,113 output tokens per index task against 4,911, and output is where the money goes
  • No longer the recommended default — OpenAI's models index now opens with Astra

Best for

  • Agents, tool use and long-horizon coding work
  • Anyone shipping to production this quarter
  • Cost-sensitive workloads that still need a frontier model
§ 06 / SCENARIOS

Picked by scenario.

More scenarios →
01

Building an agent that runs for hours

Long tool-use loops, a growing transcript, and a task that only counts if the agent finishes it.

Reasoning: Sol, and not narrowly. It scores 57.8 to Astra's 51.5 on AA's Agentic Index and 60.5 to 56.5 on GDPval, while costing 2.5x less per token. There is no reading of this workload where paying more for Astra improves the outcome — the newer model is behind on exactly the axis the work sits on.

Picked
GPT-5.6 Sol
Runner-up: Astra only if the agent's failure mode is confidently inventing facts
02

Research where a wrong answer is expensive

Legal, medical or financial analysis, where a fabricated citation costs more than the whole month's token bill.

Reasoning: Astra. AA-Omniscience rewards correct answers, penalises hallucinations and does not punish a refusal — and Astra scores 43.4 to Sol's 22.0. Humanity's Last Exam agrees at 54.7% against 49.5%. This is the one workload where the 2.5x premium is buying the thing you are actually short of.

Picked
GPT-6 Astra
Runner-up: Sol with a retrieval layer, if you can constrain it to cited sources
03

You need EU data residency and low latency

Your prompts have to be processed in-region, and you were counting on Fast mode to keep response times down.

Reasoning: Then Astra cannot do the job. OpenAI states plainly that Fast mode is unavailable for GPT-6 Astra with EU data residency and directs those requests to Standard processing; Sol carries no such restriction. The 10% regional uplift applies to both, so the surcharge is not what decides it — the missing tier is. If the work genuinely needs Astra's reasoning depth, run it Standard in-region and budget for the extra wall-clock rather than assuming Fast mode will be there.

Picked
GPT-5.6 Sol
Runner-up: Astra on Standard processing, if the latency budget can absorb it
04

Sizing next year's budget

You need a number that survives contact with finance, and the model choice follows from it.

Reasoning: Read the ratio twice. Astra is 2.5x Sol on every published rate, but only 1.49x on the measured cost of finishing Artificial Analysis's whole index, because it needs a fifth fewer output tokens. And Sol's card is promotional at least through November 21 — if it reverts to $5 and $30, the premium falls to 2x input and 1.67x output without either model changing.

Picked
GPT-5.6 Sol
Runner-up: Astra, if your work is answer-shaped rather than agent-shaped
05

Documents and screenshots, not chat

Contracts, slide decks, scanned forms — the input is as much image as text.

Reasoning: Astra, on the only vision number either model publishes here: 86.9% on MMMU-Pro against 83.4%. Both accept text and image input and emit text only, both carry a 1,050,000-token context with a 922,000-token input ceiling, and both reprice the entire request at 2x input and 1.5x output above 272K tokens — so a very large document costs double on either.

Picked
GPT-6 Astra
Runner-up: Sol, if the pages are clean text and the volume is high

Frequently asked.

Common questions about this comparison, with sources where they matter.

Q · 01 Is GPT-6 Astra actually better than GPT-5.6 Sol? +
On the composite, barely: 61.2 against 60.9 on Artificial Analysis's Intelligence Index v4.1.1, which is a smaller gap than most benchmarks move between re-runs. The components are where the answer lives, and they split cleanly. Astra wins knowledge reliability (43.4 against 22.0 on AA-Omniscience), Humanity's Last Exam (54.7% against 49.5%), GPQA Diamond and MMMU-Pro. Sol wins agentic work (57.8 against 51.5), GDPval-AA v2, τ³-Banking, SciCode and AA-LCR. So the honest answer is that OpenAI shipped a model that knows more and plans worse, and charged 2.5x for it.
Q · 02 Why is the price exactly 2.5x on everything? +
Because OpenAI scaled the whole card rather than repricing it line by line. Input $10 against $4, cached input $1 against $0.40, cache writes $12.50 against $5, output $50 against $20, batch $5 and $25 against $2 and $10 — every cell is the same multiple. That is unusual and it has a practical consequence: no workload mix escapes the premium. Normally you can shift a comparison by leaning on caching or batching, because vendors discount those unevenly; here the ratio is invariant, and our blended figure lands at exactly the same 2.5x ($6.82 against $2.73 per million on the site's 92/8 agentic mix at an 82% cache-hit rate).
Q · 03 So is Astra 2.5x more expensive in practice? +
No — about 1.49x. Artificial Analysis publishes what it spent running the whole index on each model: $3,013 on Astra against $2,017 on Sol. The rates are 2.5x apart, but Astra needs 4,911 output tokens per task against Sol's 6,113, and output is where reasoning models spend the money (of Astra's $2,105 output bill, $1,824 was reasoning tokens). Conciseness pulls the measured multiple 40% below the sticker multiple. This is the same effect described in list price versus agent cost: the rate card ranks models by the number that is not the bill.
Q · 04 Can I use GPT-6 Astra today? +
Yes — and that answer changed on September 6. Astra shipped on September 3 to Trusted Access Program enterprises, with API and paid-plan access promised in the coming days; OpenAI delivered inside three, without announcing it. The clearest tell is the model page's own rate-limit table, which now runs from Tier 1 upward — the tier any account reaches after spending $5 — where an approval-gated model would have no ladder at all. The docs also now open by pointing developers who are unsure where to start at Astra. One route remains gated, and it is Microsoft's rather than OpenAI's: Foundry sells Astra through its own Limited Access Program. Separately, if you work inside the Daybreak programme: its aliases gpt-daybreak-blue-latest and gpt-daybreak-red-latest still resolve to gpt-5.6-sol and gpt-5.6-cyber, not to Astra.
Q · 05 Will Sol's price go back up? +
OpenAI's documentation says Sol's promotional pricing is available at least through November 21, 2026, and describes the current card as a 20% cut on input and a 33% cut on output. The pre-promotion rate was $5 and $30. Two things follow. First, the $4/$20 you are budgeting against has a published expiry and no published successor — OpenAI prints only the promotional card. Second, if it does revert, Astra's premium shrinks from 2.5x on both sides to 2x on input and 1.67x on output, which changes the arithmetic without either model changing at all.
Q · 06 What is the 99.9% score I keep seeing quoted for Astra? +
It is real, and it is harness-dependent. ARC Prize ran Astra on the ARC-AGI-3 Semi-Private set twice and published both numbers: 62.7% for about $26K under ARC's own Standard harness at max effort, and 99.9% for about $19K under a Provider Adapter harness that preserves opaque reasoning state between requests and compacts long conversations so the model can reuse prior work. OpenAI's launch material quotes the 99.9%. The 37-point spread came from the scaffolding, not the model — and the run that scored higher also cost 28% less, which is the reverse of the usual effort-for-money trade.
Q · 07 Does the context window differ? +
No. Both carry a 1,050,000-token context window with a 922,000-token maximum input and 128,000 maximum output, and both reprice the entire request at 2x input and cache rates and 1.5x output once a prompt exceeds 272,000 input tokens. That threshold is the number to design around: a 300K-token prompt does not cost slightly more than a 272K one, it costs double on the whole call. Where they differ is what they do with a long context — Sol scores 77.7% on AA-LCR against Astra's 74.3%.
Q · 08 Which one should a small team pick? +
Sol, unless your product's failure mode is confident invention. It is generally available, 2.5x cheaper on every rate, faster at 76.5 output tokens a second, and better at agentic work — and its composite score is three tenths of a point behind. Astra earns its money on a narrow and important axis: knowledge reliability, where 43.4 against 22.0 is not a rounding difference. If you are unsure, the cheap test is to run your own evals on Sol now and re-run them on Astra when the gate opens, rather than buying access on the strength of a launch post.