Last verified
NEWEST GEMINI FLASH1M CONTEXTMULTIMODAL + PDF65K MAX OUTPUTHALF PRICE UNTIL 2027

Gemini 3.7 Flash API Pricing

Gemini 3.7 Flash shipped on August 14, 2026 as Google's newest Flash model, described as built for complex coding, agentic workflows and reliable multi-step execution. It arrives on the same rate card as Gemini 3.6 Flash, which Google moved onto the same schedule the same day: $0.75/M input, $3.75/M output and $0.075/M cached input through December 31, 2026, then standard pricing of $1.50 / $7.50 / $0.15 from January 1, 2027. Because the two are priced identically, the choice between them is about capability, not cost.

Input - per 1M tokens
$0.75/M
$1.50 from Jan 1, 2027 -50%
Output - per 1M tokens
$3.75/M
$7.50 from Jan 1, 2027 -50%
Cached input - per 1M tokens
$0.075/M
$0.15 from Jan 1, 2027 -90%
Effective - agentic blend
$0.481/M
92/8 split - 82% cache
§ 01 / TERMINAL

Run the numbers.

Live calculator pre-loaded with the Gemini 3.7 Flash rates in effect today — $0.75 input, $3.75 output, $0.075 cached. All three double on January 1, 2027, so a workload sized here costs twice as much from that date. Gemini 3.6 Flash prices identically, so switching between the two changes capability rather than the bill.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Estimate · gemini-tokenizer-estimate · ≈3.85 chars/token Auto-counts as you type

This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.

Characters 580
Words 90
Tokens (estimated) 151 tokens
Cost as input · uncached $0.00011 USD
Cost as output · uncached $0.00057 USD
Cost as cached input $0.00001 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Gemini 3.7 Flash Current $0.75 cache $0.075 $3.75 $0.481 newest Flash 1M Coding, agents, multi-step execution
Gemini 3.6 Flash $0.75 cache $0.075 $3.75 $0.481 same price, prior gen 1M Previous-generation Flash
Gemini 3.5 Flash $1.50 cache $0.15 $9.00 $1.08 pricier output 1M The model 3.6 replaces
Gemini 3.1 Pro Preview $2.00 cache $0.20 $12.00 $1.44 pricier 1M Google's frontier reasoning tier
Gemini 3.5 Flash-Lite $0.30 cache $0.03 $2.50 $0.272 cheaper 1M High-volume, simpler tasks
Gemini 3.1 Flash-Lite $0.25 cache $0.025 $1.50 $0.18 cheapest Gemini 1M Cheapest per-token Gemini tier
Claude Haiku 4.5 $1.00 cache $0.10 $5.00 $0.682 cheaper 200K Anthropic's fast tier
GPT-5.4 mini $0.75 cache $0.075 $4.50 $0.541 cheaper 400K OpenAI's mid tier
DeepSeek V4 Flash $0.44 cache $0.014 $1.32 $0.189 cheaper 1M Budget reasoning and coding

Frequently asked.

Gemini 3.7 Flash pricing questions, with the dated schedule kept separate from the rate in effect today.

Q · 01 How much does Gemini 3.7 Flash cost? +
Google lists gemini-3.7-flash at $0.75/M input, $3.75/M output including thinking tokens, and $0.075/M cached input. Those rates run through December 31, 2026; from January 1, 2027 the standard pricing is $1.50 / $7.50 / $0.15.
Q · 02 Is Gemini 3.7 Flash more expensive than 3.6 Flash? +
No — they are priced identically, down to the cached rate and the January step-up. Google shipped 3.7 and moved 3.6 onto the same schedule on the same day. Under our standard blend both land at $0.48/M effective, so the decision between them is a capability decision.
Q · 03 What does the January 1, 2027 change mean for my bill? +
Every rate doubles. A workload costing $1,000 a month at today's rates costs $2,000 from that date with no change on your side. Cache storage moves too, from $0.50 to $1.00 per 1M tokens per hour. We re-verify this page against Google's pricing page and will record what actually happens on the date rather than assuming the schedule holds.
Q · 04 What can Gemini 3.7 Flash take as input? +
Text, images, video, audio and PDF, returning text. The model page lists a 1,048,576-token input limit and a 65,536-token output limit, with context caching supported.
Q · 05 Are thinking tokens billed separately? +
No — Google's output price is explicitly "including thinking tokens", so reasoning is billed at the same $3.75/M as visible output. There is no separate reasoning rate to budget for, but heavy thinking still raises the output token count.