Gemini 3.8 Flash API Pricing
Gemini 3.8 Flash shipped on September 2, 2026, described by Google as its most intelligent Flash model, built for long-horizon software engineering, autonomous agents and complex enterprise workflows. It lands on the same rate card as 3.7 Flash and 3.6 Flash: $0.75/M input, $3.75/M output and $0.075/M cached input through December 31, 2026, then standard pricing of $1.50 / $7.50 / $0.15 from January 1, 2027. That is three consecutive Flash generations at one price, so upgrading costs nothing and the only question is capability.
Run the numbers.
Live calculator on Google's current Flash rate. Everything here doubles on January 1, 2027 — run the numbers twice if the workload outlives this year.
Real-world presets.
Codebase-scale migration
Reading 100-page contracts
Ticket triage
Research planning turn
Paste text. See tokens. See cost.
This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.
| Model | Input /M | Output /M | Effective blended | Context | Best for |
|---|---|---|---|---|---|
| Gemini 3.8 Flash Current | $0.75 cache $0.075 | $3.75 | $0.481 agentic 92/8 | 1M | Google's newest Flash - agents and long-horizon coding |
| Gemini 3.7 Flash | $0.75 cache $0.075 | $3.75 | $0.481 identical card, previous generation | 1M | The row this one demotes |
| Gemini 3.6 Flash | $0.75 cache $0.075 | $3.75 | $0.481 identical card, two generations back | 1M | Same price, older capability |
| Gemini 3.5 Flash | $1.50 cache $0.15 | $9.00 | $1.08 pricier output | 1M | The model 3.6 replaces |
| Gemini 3.1 Pro Preview | $2.00 cache $0.20 | $12.00 | $1.44 pricier | 1M | Google's frontier reasoning tier |
| Gemini 3.5 Flash-Lite | $0.30 cache $0.03 | $2.50 | $0.272 cheaper | 1M | High-volume, simpler tasks |
| Gemini 3.1 Flash-Lite | $0.25 cache $0.025 | $1.50 | $0.18 cheapest Gemini | 1M | Cheapest per-token Gemini tier |
| Claude Haiku 4.5 | $1.00 cache $0.10 | $5.00 | $0.682 cheaper | 200K | Anthropic's fast tier |
| GPT-5.4 mini | $0.75 cache $0.075 | $4.50 | $0.541 cheaper | 400K | OpenAI's mid tier |
| DeepSeek V4 Flash | $0.44 cache $0.014 | $1.32 | $0.189 cheaper | 1M | Budget reasoning and coding |
Frequently asked.
What the identical rate card across three Flash generations means, what changes on January 1, 2027, and what this model takes as input.
Q · 01 How much does Gemini 3.8 Flash cost? +
$0.75/M input, $3.75/M output (thinking tokens included) and $0.075/M cached input through December 31, 2026. From January 1, 2027 those become $1.50, $7.50 and $0.15. Under this site's 92/8 agentic blend at an 82% cache-hit rate the effective rate today is $0.481/M. A free tier is available with lower limits.Q · 02 Is it more expensive than 3.7 or 3.6 Flash? +
Q · 03 What changes on January 1, 2027? +
$0.75 → $1.50, output $3.75 → $7.50, cached input $0.075 → $0.15. Cache storage also doubles, from $0.50 to $1.00 per million tokens per hour. This page stores the rate in effect today; the scheduled one is recorded alongside it and will be re-checked on the day rather than assumed — a vendor announcing a change is not the same as a vendor applying it.Q · 04 What can it take as input? +
1,048,576 tokens and the output limit 65,536. Caching, code execution, function calling, file search and grounding with Google Maps are supported; computer use is in preview. Batch, Flex and Priority inference are all available.