Last verified
GEMINI 3.5 LITE GA1M CONTEXTMULTIMODALHIGH THROUGHPUTNO AUDIO SURCHARGE

Gemini 3.5 Flash-Lite API Pricing

Gemini 3.5 Flash-Lite shipped GA on July 21, 2026 for high-volume agentic tasks, translation and simple data processing. It lists $0.30/M input and $2.50/M output, with cache hits at $0.03/M. Note the rate card before switching: per token it is dearer than Gemini 3.1 Flash-Lite. Pulled directly from ai.google.dev daily.

Input - per 1M tokens
$0.30/M
Covers audio too new
Output - per 1M tokens
$2.50/M
Incl. thinking tokens new
Cached input - per 1M tokens
$0.03/M
Storage $1/M-hour -90%
Effective - agentic blend
$0.27/M
92/8 split - 82% cache
§ 01 / TERMINAL

Run the numbers.

Live calculator pre-loaded with current Gemini 3.5 Flash-Lite rates. Tweak spend, output mix, or cache assumptions and share the URL to share the calculation.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Estimate · gemini-tokenizer-estimate · ≈3.85 chars/token Auto-counts as you type

This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.

Characters 683
Words 107
Tokens (estimated) 177 tokens
Cost as input · uncached $0.00005 USD
Cost as output · uncached $0.00044 USD
Cost as cached input $0.00001 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Gemini 3.5 Flash-Lite Current $0.30 cache $0.03 $2.50 $0.27 agentic 92/8 1M High-volume agentic + translation
Gemini 3.1 Flash-Lite $0.25 cache $0.03 $1.50 $0.18 cheaper on text 1M Cheapest per-token Gemini for text
Gemini 2.5 Flash $0.30 cache $0.03 $2.50 $0.27 same rate 1M Prior-gen twin of this rate card
Gemini 3.6 Flash $1.50 cache $0.15 $7.50 $0.96 pricier 1M Step up for harder agentic work
GPT-5.4 nano $0.20 cache $0.02 $1.25 $0.15 cheaper 400K OpenAI's budget tier
DeepSeek V4 Flash $0.14 cache $0.0028 $0.28 $0.05 cheaper 1M Budget reasoning and coding

Frequently asked.

Practical pricing questions, separated from calculator assumptions and regional taxes.

Q · 01 What is Gemini 3.5 Flash-Lite priced at? +
Google's paid-tier pricing page lists $0.30/M input, $2.50/M output and $0.03/M cached input, plus $1.00 per 1M tokens per hour of explicit cache storage. The single input rate covers text, image, video and audio.
Q · 02 Google calls it the most cost-efficient GA model. Is it the cheapest Gemini? +
Not per token. Gemini 3.1 Flash-Lite is cheaper on both sides — $0.25/M input and $1.50/M output against $0.30 and $2.50 here — and Gemini 2.5 Flash carries an identical rate card. Any saving has to come from this model finishing a task in fewer tokens, which is a claim to test on your own workload, not a lower rate.
Q · 03 When is it cheaper than Gemini 3.1 Flash-Lite? +
On audio. 3.1 Flash-Lite charges $0.50/M for audio input against $0.25/M for text, while this model bills every modality at $0.30/M. Audio-heavy pipelines can therefore land cheaper here; text-only bulk work will not.
Q · 04 Does output pricing include thinking tokens? +
Yes. Google's price table labels the output row "including thinking tokens", so reasoning is billed at the same $2.50/M. Thinking is supported on this model and there is no separate reasoning rate.
Q · 05 How is the effective price calculated? +
AI//COST uses the same 92/8 agentic blend everywhere: 92% of tokens are input, 82% of which are served from cache. That gives a blended input of $0.0786/M and an effective cost of $0.27/M. Every shelf row above uses the same formula.
Q · 06 What does Batch or Flex inference cost? +
The model page marks Batch API, Flex and Priority inference as supported, but Google's pricing page publishes no per-model rate for those tiers on this model today. This page records the standard tier only rather than assuming the usual discount.
Q · 07 What context window does it support? +
1,048,576 input tokens and 65,536 output tokens. Inputs cover text, image, video, audio and PDF; output is text only. Computer use, Live API, audio generation and image generation are not supported — for computer use, see Gemini 3.6 Flash.
Q · 08 How accurate is the tokenizer estimate? +
The browser widget uses a gemini-tokenizer-estimate chars-per-token ratio for English text — Google publishes no downloadable tokenizer, so this is an estimate, not exact counting. Actual billing comes from the API's usage metadata and differs for code, non-Latin scripts and multimodal input.