Last verified

Gemini Pro vs Gemini Flash — which Gemini tier actually earns the call?

Google's tier ladder inverted. Gemini 3.6 Flash (GA, Jul 21) is cheaper ($1.50 / $7.50 vs $2 / $12), about 2× faster, and outscores the older Gemini 3.1 Pro preview on the aggregate intelligence index. Pro keeps a hard-science edge — GPQA 94.1 vs 92.8, HLE 44.7 vs 38.3. Choose Pro for deep reasoning and long-document analysis; choose Flash for volume, agents and anything latency-sensitive.

§ 01 / VERDICT

Who wins, category by category.

Skip to decision tree →
Category Winner Margin
Hard-science reasoning · GPQA Diamond AGemini Pro Pro 94.1 vs Flash 92.8 — real but narrow (BenchLM, Jul 26 2026)
Frontier reasoning · Humanity's Last Exam AGemini Pro Pro 44.7 vs Flash 38.3 — the widest gap in Pro's favour
Long-context reasoning · AA-LCR AGemini Pro Pro 72.7 vs Flash 69.7 — reasoning over long documents, not just recall
Scientific coding · AA-SciCode AGemini Pro Pro 58.9 vs Flash 52.7 — research-grade code still favours Pro
Aggregate intelligence · AA Intelligence Index BGemini Flash Flash 50.1 vs Pro 46.5 — the newer Flash outscores the older Pro preview
Agentic tool use · AA Agentic Index BGemini Flash Flash 38.7 vs Pro 21.4 — the single largest gap on this page
Agentic coding · Terminal-Bench 2.1 BGemini Flash Flash 77.5 vs Pro 73.8 — Flash was tuned for the terminal loop
General coding · AA Coding Index ·Tie 68.8 vs 69.2 — inside the noise, call it level
API list price · per 1M tokens BGemini Flash $1.50 / $7.50 vs $2 / $12 — 25% cheaper in, 38% cheaper out
Long-prompt economics · above 200K tokens BGemini Flash Pro doubles to $4 / $18 past 200K; Flash stays flat at $1.50 / $7.50
Speed · output tokens/sec BGemini Flash ~235 vs ~109 tokens/s — Flash is roughly twice as fast
Access & stability · GA vs preview BGemini Flash Flash is GA with a free API tier; 3.1 Pro is a preview with none
Best overall ·Depends See the decision tree below
CHOOSE A · GEMINI 3.1 PRO

If the job is hard thinking, not throughput.

  • Hard science — GPQA Diamond 94.1 vs 92.8 and Humanity's Last Exam 44.7 vs 38.3
  • Long-document reasoning — AA-LCR 72.7 vs 69.7, the benchmark that tests reasoning across a long context rather than needle retrieval
  • Scientific code — AA-SciCode 58.9 vs 52.7, and CritPt physics 17.7 vs 10.6
  • App parity — Google AI Pro ($19.99/mo) lifts the daily 3.1 Pro allotment, so API answers match the Pro-model output you see in the app
  • Low-volume economics — at $2 / $12 per 1M it runs about 1.5× Flash's blended cost, a premium that only stays small if you send few requests
CHOOSE B · GEMINI 3.6 FLASH

If you need volume, agents, or low latency.

  • Cheaper — $1.50 / $7.50 against $2 / $12, and no tier-2 jump when a prompt crosses 200K tokens
  • Faster — about 235 output tokens per second versus roughly 109 for Pro
  • Agentic edge — Terminal-Bench 2.1 77.5 vs 73.8 and an AA Agentic Index nearly double Pro's
  • GA, not preview — a stable slug, a free API tier, and the default model in the Gemini app since Jul 21, 2026
  • Fresher knowledge — a March 2026 training cutoff against Pro's January 2025
  • Cheaper caching — cache storage bills $1.00 per 1M tokens per hour versus Pro's $4.50
§ 02 / PRICING

What it actually costs.

Cost calculator →
Aspect Gemini Pro Gemini Flash
Gemini app accessWhich consumer plan unlocks it verified Jul 27 $19.99/mo · AI Pro Free app users get only a daily allotment of 3.1 Pro; Google AI Pro ($19.99/mo) buys 4× free limits and the 1M window. Google AI Plus is $4.99/mo for 2× limits at a 128k window $0 · app default 3.6 Flash became the Gemini app's default model on Jul 21, 2026 and is included at every tier, free through Ultra B wins
Free API tierAI Studio + Gemini API at $0 verified Jul 27 ✗ Paid only Google's rate card lists no free tier for the 3.1 Pro preview — every call bills, on standard, Batch, Flex and Priority alike ✓ Available The standard rate card lists a free tier for 3.6 Flash; Batch, Flex and Priority remain paid-only B wins
API · inputper 1M tokens · prompts ≤ 200K · from snapshot verified May 18 $2.00 $1.50 B wins
API · outputper 1M tokens · thinking tokens included · from snapshot verified May 18 $12.0 $7.50 B wins
API · cached inputper 1M cache-hit tokens · from snapshot verified May 18 $0.20 $0.15 B wins
Effective API costBlended agentic workload $/1M · from snapshot verified May 18 $1.44 $0.96 B wins
API context windowMax input tokens · from snapshot even verified May 18 1M 1M
Real cost / 1M charsTokenizer-adjusted prose — the tokenizer tax est. verified May 18 $0.52 $0.39 B wins
§ 03 / FEATURES

Feature-by-feature, side by side.

Download CSV →
Capability Gemini Pro Gemini Flash
Release status ~ Preview · May 7, 2026 ✓ GA (Stable) · Jul 21, 2026
Context window 1,048,576 tokens 1,048,576 tokens
Max output tokens 65,536 65,536
Price above 200K tokens $4 / $18 — doubles $1.50 / $7.50 — flat
Batch / Flex rate $1.00 in / $6.00 out $0.75 in / $3.75 out
Priority rate $3.60 in / $21.60 out $2.70 in / $13.50 out
Cache storage $4.50 / 1M tokens / hour $1.00 / 1M tokens / hour
Output speed ~109 tokens/sec ~235 tokens/sec
Knowledge cutoff Jan 2025 Mar 2026
Accepted inputs Text, image, video, audio, PDF Text, image, video, audio, PDF
Thinking / reasoning control
Computer use ~ Not listed on the model card ✓ Preview
Code execution
Function calling
Search + Maps grounding ✓ 5,000 prompts/mo free ✓ 5,000 prompts/mo free
Structured outputs
URL context
File search ~ AI Studio only
Context caching ✓ $0.20 / 1M cached ✓ $0.15 / 1M cached
Live API
Image / audio generation
Gemini app availability ~ Daily allotment free · full on AI Pro ✓ Default model, all tiers
§ 04 / BENCHMARKS

The numbers, not the spin.

Aggregate intelligence · AA Intelligence Index
Gemini Pro
46.5%
Gemini Flash
50.1%
BenchLM head-to-head · Artificial Analysis Intelligence Index · Gemini 3.1 Pro vs Gemini 3.6 Flash · data verified Jul 26, 2026
Reasoning · GPQA Diamond
Gemini Pro
94.1%
Gemini Flash
92.8%
BenchLM · AA-GPQA Diamond · Gemini 3.1 Pro vs Gemini 3.6 Flash · Jul 26, 2026 · a saturated benchmark — treat 1.3 points as narrow
Frontier reasoning · Humanity's Last Exam
Gemini Pro
44.7%
Gemini Flash
38.3%
BenchLM · AA-HLE · Gemini 3.1 Pro vs Gemini 3.6 Flash · Jul 26, 2026 · Pro's clearest remaining lead
Agentic coding · Terminal-Bench 2.1
Gemini Pro
73.8%
Gemini Flash
77.5%
BenchLM · aaTerminalBench21 · Gemini 3.1 Pro vs Gemini 3.6 Flash · Jul 26, 2026 · Google's own launch card reports 73.8 vs 78.0 on the same test
Coding · AA Coding Index
Gemini Pro
68.8%
Gemini Flash
69.2%
BenchLM · Artificial Analysis Coding Index · Gemini 3.1 Pro vs Gemini 3.6 Flash · Jul 26, 2026 · 0.4 points apart — a tie
§ 05 / DEEP DIVE

What each does best.

Brand hubs →
A · GOOGLE — GEMINI 3.1 PRO PREVIEW

Gemini Pro

Google's deep-thinking tier, still in preview after two and a half months — ahead on hard-science and long-document reasoning, behind on price, speed and agentic work.

Strengths

  • Hard science — the best Gemini on GPQA Diamond (94.1) and Humanity's Last Exam (44.7)
  • Long-context reasoning — leads AA-LCR 72.7 vs 69.7, which tests reasoning across a filled window rather than retrieval from it
  • Research code — AA-SciCode 58.9 vs 52.7 and CritPt physics 17.7 vs 10.6
  • Full 1M window — the same 1,048,576-token input limit as Flash, with 65,536 tokens of output
  • App parity — the model Google AI Pro and Ultra subscribers get, so API and app behaviour line up

Weaknesses

  • Still a preview slug after launching on May 7, 2026 — no GA guarantee, and Google shipped three new models in July without a 3.5 or 3.6 Pro
  • No free API tier at all, on any inference mode
  • Price doubles to $4 / $18 per 1M once a prompt crosses 200K tokens — exactly where a 1M window gets used
  • Slow to first token — roughly 31s versus 12.8s for Flash on Artificial Analysis's standard workload
  • Cache storage costs $4.50 per 1M tokens per hour, 4.5× Flash's rate
  • January 2025 knowledge cutoff, a year and two months behind Flash

Best for

  • Scientific and technical questions where one wrong answer costs more than a thousand API calls
  • Reasoning over a long report, contract or codebase in a single pass
  • Research-grade code and quantitative work
  • Matching what a Google AI Pro subscriber sees in the app
B · GOOGLE — GEMINI 3.6 FLASH

Gemini Flash

The volume workhorse that overtook its own Pro tier — GA since July 21, cheaper, twice as fast, and ahead on every agentic measure.

Strengths

  • Cheaper on every mode — $1.50 / $7.50 standard, $0.75 / $3.75 on Batch and Flex, $2.70 / $13.50 on Priority
  • Flat above 200K — no tier-2 rate, so long prompts cost what short ones do per token
  • Agentic lead — Terminal-Bench 2.1 77.5 vs 73.8 and an AA Agentic Index of 38.7 against Pro's 21.4
  • Speed — about 235 output tokens per second, roughly double Pro's ~109
  • GA and free-tier eligible — a stable slug with a free tier in AI Studio and the Gemini API
  • Computer use — listed in preview on the model card, which the 3.1 Pro card does not carry
  • Fresher — March 2026 knowledge cutoff

Weaknesses

  • Loses the hard-science benchmarks — GPQA 92.8 vs 94.1, HLE 38.3 vs 44.7
  • Weaker on long-context reasoning (AA-LCR 69.7) and scientific code (AA-SciCode 52.7)
  • Named "Flash" but priced like a mid-tier model — Flash-Lite at $0.30 / $2.50 is the actual cheap tier
  • Time to first token is still around 12.8s, so it is fast at streaming, not instant to start

Best for

  • High-volume classification, extraction and summarisation pipelines
  • Agentic coding loops and tool-calling agents
  • Latency-sensitive chat and product surfaces
  • Anyone starting on the free API tier before committing budget
§ 06 / SCENARIOS

Picked by scenario.

More scenarios →
01

Running a million documents through a classifier

You tag, extract and summarise a large corpus nightly. Each item is short, the volume is enormous, and quality only has to clear a fixed bar.

Reasoning: Flash is cheaper on every inference mode, twice as fast, and the quality gap on routine extraction does not show up in the benchmarks that matter here. Run it on Batch at $0.75 / $3.75 and the bill halves again. Test Flash-Lite at $0.30 / $2.50 before you settle — on simple labelling it often clears the same bar for a fifth of the output cost.

Picked
Gemini Flash
Runner-up: Gemini 3.5 Flash-Lite where the task is simple enough
02

Analyst working through hard technical literature

You read graduate-level science and engineering material and need answers you can defend, not answers you have to re-check.

Reasoning: This is the one place Pro's lead is unambiguous: GPQA Diamond 94.1 vs 92.8, Humanity's Last Exam 44.7 vs 38.3, AA-SciCode 58.9 vs 52.7. The volumes are low, so the 33% input premium is noise against the cost of a wrong answer. Pro's January 2025 cutoff is the catch — pair it with Search grounding for anything post-2025.

Picked
Gemini Pro
Runner-up: Flash with grounding enabled for anything time-sensitive
03

Engineer running an agentic coding loop

Your agent edits files, runs tests and iterates in a terminal. Each run burns tens of thousands of tokens and every extra second of latency compounds across the loop.

Reasoning: Flash wins the tests that model this directly — Terminal-Bench 2.1 77.5 vs 73.8, and an AA Agentic Index of 38.7 against 21.4. It also streams at roughly twice the speed, which matters more in a loop than in a single answer. Pro's reasoning edge does not convert into completed terminal tasks here.

Picked
Gemini Flash
Runner-up: Pro for the one hard architectural question mid-session
04

Reviewing a 400K-token contract set in one pass

You load an entire deal room into the window and ask cross-referencing questions that require holding the whole thing at once.

Reasoning: Both models take 1,048,576 tokens, but Pro leads AA-LCR 72.7 vs 69.7 — the test for reasoning across a filled window rather than retrieving from it. Budget for the tier-2 rate: above 200K tokens Pro bills $4 / $18 per 1M while Flash stays at $1.50 / $7.50, so a single long pass on Pro costs roughly 2.7× the Flash equivalent. Worth it for a one-off review, not for a recurring pipeline.

Picked
Gemini Pro
Runner-up: Flash when the same long prompt runs daily
05

Picking a Google AI plan for everyday use

You use Gemini in the browser and on your phone and want to know whether the $19.99 subscription is doing anything for you.

Reasoning: The free tier now defaults to 3.6 Flash — the model that wins most of this page — plus a daily allotment of 3.1 Pro. Google AI Plus at $4.99/mo doubles the limits and lifts the window to 128k; Google AI Pro at $19.99/mo buys 4× the free limits, the 1M window and 5 TB of storage. Start free, and upgrade only when you actually hit the ceiling on Pro-model prompts.

Picked
Gemini Flash
Runner-up: Google AI Pro at $19.99/mo once you exhaust the daily Pro allotment

Frequently asked.

Common questions about this comparison, with sources where they matter.

Q · 01 Which Gemini should I call — Pro or Flash? +
Default to Flash and escalate to Pro only for hard reasoning. As of July 2026 the current Flash (Gemini 3.6, GA on Jul 21) is cheaper ($1.50 / $7.50 vs $2 / $12), roughly twice as fast, and ahead of the current Pro (Gemini 3.1 Pro Preview, May 7) on the aggregate Artificial Analysis Intelligence Index, 50.1 vs 46.5. Pro still leads GPQA Diamond, Humanity's Last Exam, long-context reasoning and scientific code. That is a narrower Pro remit than the naming implies.
Q · 02 Why does Flash beat Pro on so many benchmarks? +
Release dates, not tier design. Gemini 3.1 Pro Preview shipped on May 7, 2026; Gemini 3.6 Flash shipped on July 21, 2026 — two and a half months and two model generations later. Google released three models that day and none of them was a new Pro. Until a 3.5 or 3.6 Pro lands, the newest Flash is measured against a Pro that predates it, and on aggregate and agentic scores it wins.
Q · 03 How much cheaper is Flash, really? +
Less than the name suggests. List rates are $1.50 / $7.50 against $2 / $12 — 25% cheaper on input, 38% on output. On our blended agentic mix (92% input, 8% output, 82% cache hits) that works out to roughly $0.96 vs $1.44 per 1M tokens, a 1.5× gap rather than the 10× the Flash label evokes. If you want a genuinely cheap tier, Gemini 3.5 Flash-Lite is $0.30 / $2.50. Model your own split in the LLM API cost calculator.
Q · 04 Do both have the same context window? +
Yes — both accept 1,048,576 input tokens and emit up to 65,536. The difference is what it costs to fill. Pro is tier-priced: above 200K tokens per prompt the rate doubles to $4 / $18 per 1M. Flash has no tier-2 rate. So a 400K-token prompt costs about 2.7× more on Pro than on Flash, and the gap widens the more of the window you actually use.
Q · 05 Is it safe to build on Gemini 3.1 Pro? +
It is a preview model, and it has been one since May 2026. Preview slugs at Google get replaced without the deprecation notice a GA model carries — gemini-3-pro-preview was retired in March 2026. Pro also has no free API tier on any inference mode, so there is no zero-cost way to test it. If you need a stable contract, pin gemini-3.6-flash and treat the Pro preview as an escalation path rather than a foundation.
Q · 06 Which Google AI plan gives me Pro? +
The free Gemini app already runs 3.6 Flash as its default and grants a daily allotment of 3.1 Pro. Google AI Plus is $4.99/mo for 2× the free usage limits, a 128k context window and 400 GB of storage — Google cut it from $7.99 and doubled the storage in June 2026. Google AI Pro is $19.99/mo for 4× free limits, the 1M window and 5 TB. Google AI Ultra starts at $99.99/mo and is the only tier with Deep Think. See Gemini Free vs Pro for the full plan split.
Q · 07 Which one is faster? +
Flash, by about 2×. Artificial Analysis measures roughly 235 output tokens per second for Gemini 3.6 Flash against about 109 for Gemini 3.1 Pro. Time to first token favours Flash too — around 12.8s versus 31.2s on the same standard workload. Both are thinking models, so neither is instant; the difference shows up most in agent loops, where the same wait is paid on every step.
Q · 08 Can I route between them? +
That is the sensible setup. Run 3.6 Flash as the default, and escalate to 3.1 Pro only when a request trips a hard-reasoning heuristic — scientific content, multi-hop analysis over a long document, anything where a wrong answer is expensive. Because both models share a 1M window and the same input formats, the switch is a model-string change, and prompts port without a rewrite. Cache hits do not carry across models, so plan for a cold cache on every escalation.