Gemini Pro vs Gemini Flash — which Gemini tier actually earns the call?
Google's tier ladder inverted. Gemini 3.6 Flash (GA, Jul 21) is cheaper ($1.50 / $7.50 vs $2 / $12), about 2× faster, and outscores the older Gemini 3.1 Pro preview on the aggregate intelligence index. Pro keeps a hard-science edge — GPQA 94.1 vs 92.8, HLE 44.7 vs 38.3. Choose Pro for deep reasoning and long-document analysis; choose Flash for volume, agents and anything latency-sensitive.
| Category | Winner | Margin |
|---|---|---|
| Hard-science reasoning · GPQA Diamond | AGemini Pro | Pro 94.1 vs Flash 92.8 — real but narrow (BenchLM, Jul 26 2026) |
| Frontier reasoning · Humanity's Last Exam | AGemini Pro | Pro 44.7 vs Flash 38.3 — the widest gap in Pro's favour |
| Long-context reasoning · AA-LCR | AGemini Pro | Pro 72.7 vs Flash 69.7 — reasoning over long documents, not just recall |
| Scientific coding · AA-SciCode | AGemini Pro | Pro 58.9 vs Flash 52.7 — research-grade code still favours Pro |
| Aggregate intelligence · AA Intelligence Index | BGemini Flash | Flash 50.1 vs Pro 46.5 — the newer Flash outscores the older Pro preview |
| Agentic tool use · AA Agentic Index | BGemini Flash | Flash 38.7 vs Pro 21.4 — the single largest gap on this page |
| Agentic coding · Terminal-Bench 2.1 | BGemini Flash | Flash 77.5 vs Pro 73.8 — Flash was tuned for the terminal loop |
| General coding · AA Coding Index | ·Tie | 68.8 vs 69.2 — inside the noise, call it level |
| API list price · per 1M tokens | BGemini Flash | $1.50 / $7.50 vs $2 / $12 — 25% cheaper in, 38% cheaper out |
| Long-prompt economics · above 200K tokens | BGemini Flash | Pro doubles to $4 / $18 past 200K; Flash stays flat at $1.50 / $7.50 |
| Speed · output tokens/sec | BGemini Flash | ~235 vs ~109 tokens/s — Flash is roughly twice as fast |
| Access & stability · GA vs preview | BGemini Flash | Flash is GA with a free API tier; 3.1 Pro is a preview with none |
| Best overall | ·Depends | See the decision tree below |
If the job is hard thinking, not throughput.
- Hard science — GPQA Diamond 94.1 vs 92.8 and Humanity's Last Exam 44.7 vs 38.3
- Long-document reasoning — AA-LCR 72.7 vs 69.7, the benchmark that tests reasoning across a long context rather than needle retrieval
- Scientific code — AA-SciCode 58.9 vs 52.7, and CritPt physics 17.7 vs 10.6
- App parity — Google AI Pro ($19.99/mo) lifts the daily 3.1 Pro allotment, so API answers match the Pro-model output you see in the app
- Low-volume economics — at $2 / $12 per 1M it runs about 1.5× Flash's blended cost, a premium that only stays small if you send few requests
If you need volume, agents, or low latency.
- Cheaper — $1.50 / $7.50 against $2 / $12, and no tier-2 jump when a prompt crosses 200K tokens
- Faster — about 235 output tokens per second versus roughly 109 for Pro
- Agentic edge — Terminal-Bench 2.1 77.5 vs 73.8 and an AA Agentic Index nearly double Pro's
- GA, not preview — a stable slug, a free API tier, and the default model in the Gemini app since Jul 21, 2026
- Fresher knowledge — a March 2026 training cutoff against Pro's January 2025
- Cheaper caching — cache storage bills $1.00 per 1M tokens per hour versus Pro's $4.50
| Aspect | Gemini Pro | Gemini Flash |
|---|---|---|
| Gemini app accessWhich consumer plan unlocks it | $19.99/mo · AI Pro Free app users get only a daily allotment of 3.1 Pro; Google AI Pro ($19.99/mo) buys 4× free limits and the 1M window. Google AI Plus is $4.99/mo for 2× limits at a 128k window | $0 · app default 3.6 Flash became the Gemini app's default model on Jul 21, 2026 and is included at every tier, free through Ultra B wins |
| Free API tierAI Studio + Gemini API at $0 | ✗ Paid only Google's rate card lists no free tier for the 3.1 Pro preview — every call bills, on standard, Batch, Flex and Priority alike | ✓ Available The standard rate card lists a free tier for 3.6 Flash; Batch, Flex and Priority remain paid-only B wins |
| API · inputper 1M tokens · prompts ≤ 200K · from snapshot | $2.00 | $1.50 B wins |
| API · outputper 1M tokens · thinking tokens included · from snapshot | $12.0 | $7.50 B wins |
| API · cached inputper 1M cache-hit tokens · from snapshot | $0.20 | $0.15 B wins |
| Effective API costBlended agentic workload $/1M · from snapshot | $1.44 | $0.96 B wins |
| API context windowMax input tokens · from snapshot | 1M | 1M |
| Real cost / 1M charsTokenizer-adjusted prose — the tokenizer tax | $0.52 | $0.39 B wins |
| Capability | Gemini Pro | Gemini Flash |
|---|---|---|
| Release status | ~ Preview · May 7, 2026 | ✓ GA (Stable) · Jul 21, 2026 |
| Context window | 1,048,576 tokens | 1,048,576 tokens |
| Max output tokens | 65,536 | 65,536 |
| Price above 200K tokens | $4 / $18 — doubles | $1.50 / $7.50 — flat |
| Batch / Flex rate | $1.00 in / $6.00 out | $0.75 in / $3.75 out |
| Priority rate | $3.60 in / $21.60 out | $2.70 in / $13.50 out |
| Cache storage | $4.50 / 1M tokens / hour | $1.00 / 1M tokens / hour |
| Output speed | ~109 tokens/sec | ~235 tokens/sec |
| Knowledge cutoff | Jan 2025 | Mar 2026 |
| Accepted inputs | Text, image, video, audio, PDF | Text, image, video, audio, PDF |
| Thinking / reasoning control | ✓ | ✓ |
| Computer use | ~ Not listed on the model card | ✓ Preview |
| Code execution | ✓ | ✓ |
| Function calling | ✓ | ✓ |
| Search + Maps grounding | ✓ 5,000 prompts/mo free | ✓ 5,000 prompts/mo free |
| Structured outputs | ✓ | ✓ |
| URL context | ✓ | ✓ |
| File search | ~ AI Studio only | ✓ |
| Context caching | ✓ $0.20 / 1M cached | ✓ $0.15 / 1M cached |
| Live API | ✗ | ✗ |
| Image / audio generation | ✗ | ✗ |
| Gemini app availability | ~ Daily allotment free · full on AI Pro | ✓ Default model, all tiers |
The numbers, not the spin.
Gemini Pro
Google's deep-thinking tier, still in preview after two and a half months — ahead on hard-science and long-document reasoning, behind on price, speed and agentic work.
Strengths
- Hard science — the best Gemini on GPQA Diamond (94.1) and Humanity's Last Exam (44.7)
- Long-context reasoning — leads AA-LCR 72.7 vs 69.7, which tests reasoning across a filled window rather than retrieval from it
- Research code — AA-SciCode 58.9 vs 52.7 and CritPt physics 17.7 vs 10.6
- Full 1M window — the same 1,048,576-token input limit as Flash, with 65,536 tokens of output
- App parity — the model Google AI Pro and Ultra subscribers get, so API and app behaviour line up
Weaknesses
- Still a preview slug after launching on May 7, 2026 — no GA guarantee, and Google shipped three new models in July without a 3.5 or 3.6 Pro
- No free API tier at all, on any inference mode
- Price doubles to $4 / $18 per 1M once a prompt crosses 200K tokens — exactly where a 1M window gets used
- Slow to first token — roughly 31s versus 12.8s for Flash on Artificial Analysis's standard workload
- Cache storage costs $4.50 per 1M tokens per hour, 4.5× Flash's rate
- January 2025 knowledge cutoff, a year and two months behind Flash
Best for
- Scientific and technical questions where one wrong answer costs more than a thousand API calls
- Reasoning over a long report, contract or codebase in a single pass
- Research-grade code and quantitative work
- Matching what a Google AI Pro subscriber sees in the app
Gemini Flash
The volume workhorse that overtook its own Pro tier — GA since July 21, cheaper, twice as fast, and ahead on every agentic measure.
Strengths
- Cheaper on every mode — $1.50 / $7.50 standard, $0.75 / $3.75 on Batch and Flex, $2.70 / $13.50 on Priority
- Flat above 200K — no tier-2 rate, so long prompts cost what short ones do per token
- Agentic lead — Terminal-Bench 2.1 77.5 vs 73.8 and an AA Agentic Index of 38.7 against Pro's 21.4
- Speed — about 235 output tokens per second, roughly double Pro's ~109
- GA and free-tier eligible — a stable slug with a free tier in AI Studio and the Gemini API
- Computer use — listed in preview on the model card, which the 3.1 Pro card does not carry
- Fresher — March 2026 knowledge cutoff
Weaknesses
- Loses the hard-science benchmarks — GPQA 92.8 vs 94.1, HLE 38.3 vs 44.7
- Weaker on long-context reasoning (AA-LCR 69.7) and scientific code (AA-SciCode 52.7)
- Named "Flash" but priced like a mid-tier model — Flash-Lite at $0.30 / $2.50 is the actual cheap tier
- Time to first token is still around 12.8s, so it is fast at streaming, not instant to start
Best for
- High-volume classification, extraction and summarisation pipelines
- Agentic coding loops and tool-calling agents
- Latency-sensitive chat and product surfaces
- Anyone starting on the free API tier before committing budget
Running a million documents through a classifier
You tag, extract and summarise a large corpus nightly. Each item is short, the volume is enormous, and quality only has to clear a fixed bar.
Reasoning: Flash is cheaper on every inference mode, twice as fast, and the quality gap on routine extraction does not show up in the benchmarks that matter here. Run it on Batch at $0.75 / $3.75 and the bill halves again. Test Flash-Lite at $0.30 / $2.50 before you settle — on simple labelling it often clears the same bar for a fifth of the output cost.
Analyst working through hard technical literature
You read graduate-level science and engineering material and need answers you can defend, not answers you have to re-check.
Reasoning: This is the one place Pro's lead is unambiguous: GPQA Diamond 94.1 vs 92.8, Humanity's Last Exam 44.7 vs 38.3, AA-SciCode 58.9 vs 52.7. The volumes are low, so the 33% input premium is noise against the cost of a wrong answer. Pro's January 2025 cutoff is the catch — pair it with Search grounding for anything post-2025.
Engineer running an agentic coding loop
Your agent edits files, runs tests and iterates in a terminal. Each run burns tens of thousands of tokens and every extra second of latency compounds across the loop.
Reasoning: Flash wins the tests that model this directly — Terminal-Bench 2.1 77.5 vs 73.8, and an AA Agentic Index of 38.7 against 21.4. It also streams at roughly twice the speed, which matters more in a loop than in a single answer. Pro's reasoning edge does not convert into completed terminal tasks here.
Reviewing a 400K-token contract set in one pass
You load an entire deal room into the window and ask cross-referencing questions that require holding the whole thing at once.
Reasoning: Both models take 1,048,576 tokens, but Pro leads AA-LCR 72.7 vs 69.7 — the test for reasoning across a filled window rather than retrieving from it. Budget for the tier-2 rate: above 200K tokens Pro bills $4 / $18 per 1M while Flash stays at $1.50 / $7.50, so a single long pass on Pro costs roughly 2.7× the Flash equivalent. Worth it for a one-off review, not for a recurring pipeline.
Picking a Google AI plan for everyday use
You use Gemini in the browser and on your phone and want to know whether the $19.99 subscription is doing anything for you.
Reasoning: The free tier now defaults to 3.6 Flash — the model that wins most of this page — plus a daily allotment of 3.1 Pro. Google AI Plus at $4.99/mo doubles the limits and lifts the window to 128k; Google AI Pro at $19.99/mo buys 4× the free limits, the 1M window and 5 TB of storage. Start free, and upgrade only when you actually hit the ceiling on Pro-model prompts.
Frequently asked.
Common questions about this comparison, with sources where they matter.
Q · 01 Which Gemini should I call — Pro or Flash? +
$1.50 / $7.50 vs $2 / $12), roughly twice as fast, and ahead of the current Pro (Gemini 3.1 Pro Preview, May 7) on the aggregate Artificial Analysis Intelligence Index, 50.1 vs 46.5. Pro still leads GPQA Diamond, Humanity's Last Exam, long-context reasoning and scientific code. That is a narrower Pro remit than the naming implies.Q · 02 Why does Flash beat Pro on so many benchmarks? +
Q · 03 How much cheaper is Flash, really? +
$1.50 / $7.50 against $2 / $12 — 25% cheaper on input, 38% on output. On our blended agentic mix (92% input, 8% output, 82% cache hits) that works out to roughly $0.96 vs $1.44 per 1M tokens, a 1.5× gap rather than the 10× the Flash label evokes. If you want a genuinely cheap tier, Gemini 3.5 Flash-Lite is $0.30 / $2.50. Model your own split in the LLM API cost calculator.Q · 04 Do both have the same context window? +
1,048,576 input tokens and emit up to 65,536. The difference is what it costs to fill. Pro is tier-priced: above 200K tokens per prompt the rate doubles to $4 / $18 per 1M. Flash has no tier-2 rate. So a 400K-token prompt costs about 2.7× more on Pro than on Flash, and the gap widens the more of the window you actually use.Q · 05 Is it safe to build on Gemini 3.1 Pro? +
gemini-3-pro-preview was retired in March 2026. Pro also has no free API tier on any inference mode, so there is no zero-cost way to test it. If you need a stable contract, pin gemini-3.6-flash and treat the Pro preview as an escalation path rather than a foundation.Q · 06 Which Google AI plan gives me Pro? +
Q · 07 Which one is faster? +
235 output tokens per second for Gemini 3.6 Flash against about 109 for Gemini 3.1 Pro. Time to first token favours Flash too — around 12.8s versus 31.2s on the same standard workload. Both are thinking models, so neither is instant; the difference shows up most in agent loops, where the same wait is paid on every step.