Last verified
NEWEST GEMINI FLASH1M CONTEXTMULTIMODAL + PDF65K MAX OUTPUTHALF PRICE UNTIL 2027

Gemini 3.8 Flash API Pricing

Gemini 3.8 Flash shipped on September 2, 2026, described by Google as its most intelligent Flash model, built for long-horizon software engineering, autonomous agents and complex enterprise workflows. It lands on the same rate card as 3.7 Flash and 3.6 Flash: $0.75/M input, $3.75/M output and $0.075/M cached input through December 31, 2026, then standard pricing of $1.50 / $7.50 / $0.15 from January 1, 2027. That is three consecutive Flash generations at one price, so upgrading costs nothing and the only question is capability.

Input - per 1M tokens
$0.75/M
$1.50 from Jan 1, 2027 -50%
Output - per 1M tokens
$3.75/M
$7.50 from Jan 1, 2027 -50%
Cached input - per 1M tokens
$0.075/M
$0.15 from Jan 1, 2027 -90%
Effective - agentic blend
$0.481/M
92/8 split - 82% cache
§ 01 / TERMINAL

Run the numbers.

Live calculator on Google's current Flash rate. Everything here doubles on January 1, 2027 — run the numbers twice if the workload outlives this year.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Estimate · gemini-tokenizer-estimate · ≈3.85 chars/token Auto-counts as you type

This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.

Characters 841
Words 128
Tokens (estimated) 218 tokens
Cost as input · uncached $0.00016 USD
Cost as output · uncached $0.00082 USD
Cost as cached input $0.00002 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Gemini 3.8 Flash Current $0.75 cache $0.075 $3.75 $0.481 agentic 92/8 1M Google's newest Flash - agents and long-horizon coding
Gemini 3.7 Flash $0.75 cache $0.075 $3.75 $0.481 identical card, previous generation 1M The row this one demotes
Gemini 3.6 Flash $0.75 cache $0.075 $3.75 $0.481 identical card, two generations back 1M Same price, older capability
Gemini 3.5 Flash $1.50 cache $0.15 $9.00 $1.08 pricier output 1M The model 3.6 replaces
Gemini 3.1 Pro Preview $2.00 cache $0.20 $12.00 $1.44 pricier 1M Google's frontier reasoning tier
Gemini 3.5 Flash-Lite $0.30 cache $0.03 $2.50 $0.272 cheaper 1M High-volume, simpler tasks
Gemini 3.1 Flash-Lite $0.25 cache $0.025 $1.50 $0.18 cheapest Gemini 1M Cheapest per-token Gemini tier
Claude Haiku 4.5 $1.00 cache $0.10 $5.00 $0.682 cheaper 200K Anthropic's fast tier
GPT-5.4 mini $0.75 cache $0.075 $4.50 $0.541 cheaper 400K OpenAI's mid tier
DeepSeek V4 Flash $0.44 cache $0.014 $1.32 $0.189 cheaper 1M Budget reasoning and coding

Frequently asked.

What the identical rate card across three Flash generations means, what changes on January 1, 2027, and what this model takes as input.

Q · 01 How much does Gemini 3.8 Flash cost? +
Google lists $0.75/M input, $3.75/M output (thinking tokens included) and $0.075/M cached input through December 31, 2026. From January 1, 2027 those become $1.50, $7.50 and $0.15. Under this site's 92/8 agentic blend at an 82% cache-hit rate the effective rate today is $0.481/M. A free tier is available with lower limits.
Q · 02 Is it more expensive than 3.7 or 3.6 Flash? +
No — the three are priced identically, to the cent, on the same dated schedule. Google has now put three consecutive Flash generations on one rate card, so an upgrade costs nothing and the decision is purely about capability. Google positions 3.8 Flash for long-horizon software engineering and autonomous agents, and calls 3.7 Flash its previous generation as of this release.
Q · 03 What changes on January 1, 2027? +
Every token rate doubles: input $0.75 → $1.50, output $3.75 → $7.50, cached input $0.075 → $0.15. Cache storage also doubles, from $0.50 to $1.00 per million tokens per hour. This page stores the rate in effect today; the scheduled one is recorded alongside it and will be re-checked on the day rather than assumed — a vendor announcing a change is not the same as a vendor applying it.
Q · 04 What can it take as input? +
Text, image, video, audio and PDF in; text out. The input limit is 1,048,576 tokens and the output limit 65,536. Caching, code execution, function calling, file search and grounding with Google Maps are supported; computer use is in preview. Batch, Flex and Priority inference are all available.
Q · 05 Are thinking tokens billed separately? +
No. Google states the output price includes thinking tokens, so a reasoning-heavy call shows up as ordinary output on the bill. That matters when comparing against vendors that meter a separate reasoning rate.