Last verified
GEMINI 3.6 GA1M CONTEXTMULTIMODALCOMPUTER USE17% CHEAPER OUTPUT

Gemini 3.6 Flash API Pricing

Gemini 3.6 Flash is Google's new mainstream model, shipped GA on July 21, 2026. It holds Gemini 3.5 Flash's $1.50/M input but drops output to $7.50/M from $9.00 — a rare case of a vendor cutting the rate on a newer model. Cached input stays $0.15/M. Pulled directly from ai.google.dev daily.

Input - per 1M tokens
$1.50/M
Same as 3.5 Flash flat
Output - per 1M tokens
$7.50/M
Was $9.00 on 3.5 -17%
Cached input - per 1M tokens
$0.15/M
Storage $1/M-hour -90%
Effective - agentic blend
$0.96/M
92/8 split - 82% cache
§ 01 / TERMINAL

Run the numbers.

Live calculator pre-loaded with current Gemini 3.6 Flash rates. Tweak spend, output mix, or cache assumptions and share the URL to share the calculation.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Estimate · gemini-tokenizer-estimate · ≈3.85 chars/token Auto-counts as you type

This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.

Characters 626
Words 100
Tokens (estimated) 163 tokens
Cost as input · uncached $0.00024 USD
Cost as output · uncached $0.00122 USD
Cost as cached input $0.00002 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Gemini 3.6 Flash Current $1.50 cache $0.15 $7.50 $0.96 agentic 92/8 1M Mainstream agentic + multimodal work
Gemini 3.5 Flash $1.50 cache $0.15 $9.00 $1.08 pricier output 1M The model 3.6 replaces
Gemini 3.1 Pro Preview $2.00 cache $0.20 $12.00 $1.44 pricier 1M Google's frontier reasoning tier
Gemini 3.5 Flash-Lite $0.30 cache $0.03 $2.50 $0.27 cheaper 1M High-volume, simpler tasks
Gemini 3.1 Flash-Lite $0.25 cache $0.03 $1.50 $0.18 cheapest Gemini 1M Cheapest per-token Gemini tier
Claude Haiku 4.5 $1.00 cache $0.10 $5.00 $0.68 cheaper 200K Anthropic's fast tier
GPT-5.4 mini $0.75 cache $0.07 $4.50 $0.54 cheaper 400K OpenAI's mid tier
DeepSeek V4 Flash $0.14 cache $0.0028 $0.28 $0.05 cheaper 1M Budget reasoning and coding

Frequently asked.

Practical pricing questions, separated from calculator assumptions and regional taxes.

Q · 01 What is Gemini 3.6 Flash priced at? +
Google's paid-tier pricing page lists $1.50/M input, $7.50/M output and $0.15/M cached input. Explicit cache storage is billed on top at $1.00 per 1M tokens per hour.
Q · 02 Is it cheaper than Gemini 3.5 Flash? +
Yes, on output. Input is identical at $1.50/M, but output drops from $9.00/M to $7.50/M — 17% less. On our agentic blend that is $0.96/M effective against $1.08/M for 3.5 Flash. Google also says 3.6 consumes fewer output tokens per task, which would compound the saving, but that is a vendor claim we have not measured.
Q · 03 Does output pricing include thinking tokens? +
Yes. Google's price table labels the output row "including thinking tokens", so reasoning is billed at the same $7.50/M. There is no separate reasoning rate and no cheaper non-thinking variant.
Q · 04 How is the effective price calculated? +
AI//COST uses the same 92/8 agentic blend everywhere: 92% of tokens are input, 82% of which are served from cache. For Gemini 3.6 Flash that gives a blended input of $0.393/M and an effective cost of $0.96/M. Every model in the shelf above uses the identical formula, so the numbers compare directly.
Q · 05 What does Batch or Flex inference cost? +
The model page marks Batch API, Flex and Priority inference as supported, but Google's pricing page publishes no per-model rate for those tiers on this model today. We record only what is published, so this page carries the standard tier alone — check the Batch API docs before assuming the usual discount applies.
Q · 06 Is Search grounding included in the token price? +
No. Grounding with Google Search and with Google Maps is billed separately: 5,000 prompts per month free, shared across the Gemini 3 family, then $14 per 1,000 queries. One request can trigger several queries, and each is charged.
Q · 07 What context window does it support? +
1,048,576 input tokens and 65,536 output tokens. Inputs cover text, image, video, audio and PDF; output is text only. Computer use is supported in preview, while Live API, audio generation and image generation are not.
Q · 08 How accurate is the tokenizer estimate? +
The browser widget uses a gemini-tokenizer-estimate chars-per-token ratio for English text — Google publishes no downloadable tokenizer, so this is an estimate, not exact counting. Actual billing comes from the API's usage metadata and differs for code, non-Latin scripts and multimodal input.