Last verified
VERIFIED PRICEPEAK / OFF-PEAK BILLING1M CONTEXTTHINKING SAME PRICECACHE HIT 31x CHEAPER

DeepSeek V4 Flash API Pricing

DeepSeek bills by time of day from 16:00 UTC on August 16, 2026. Peak hours are 01:00-04:00 and 06:00-10:00 UTC; every other hour is off-peak at exactly half the peak rate. The tiles below carry the peak rate, because DeepSeek describes off-peak as "half the peak rates" - peak is the list price and off-peak is the discount. At peak this model is $0.44/M input, $1.32/M output and $0.014/M cached input; off-peak it is $0.22 / $0.66 / $0.007. That replaced a flat $0.14 / $0.28 / $0.0028 card, so the effective blended rate roughly quadrupled at peak and doubled off-peak. A cache hit still costs about one thirty-first of a miss. Thinking and non-thinking modes bill from the same rate card, and thinking is the default. Since July 31, 2026 the deepseek-v4-flash id serves checkpoint DeepSeek-V4-Flash-0731; the id and the prices did not change with it.

Input - per 1M tokens
$0.44/M
Off-peak $0.22/M peak
Output - per 1M tokens
$1.32/M
Off-peak $0.66/M peak
Cached input
$0.014/M
Off-peak $0.007/M peak
Effective - agentic blend
$0.189/M
92/8 split - 82% cache
§ 01 / TERMINAL

Run the numbers.

Calculator pre-loaded with DeepSeek V4 Flash PEAK rates — halve every result for an off-peak run, which covers 17 of every 24 hours. Tweak spend, output mix, or cache hit rate to compare this model with nearby alternatives.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Calibrated · measured on the vendor's tokenizer · 2026-06-10 Auto-counts as you type

Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (deepseek-ai/DeepSeek-V3.2-Exp, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.

Characters 444
Words 79
Tokens (estimated) 84 tokens
Cost as input · uncached $0.00004 USD
Cost as output · uncached $0.00011 USD
Cost as cached input $0.000001 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
DeepSeek V4 Flash $0.44 cache $0.014 $1.32 $0.189 current page - peak 1M Cheap 1M-context agent loops
DeepSeek V4 Pro $1.32 cache $0.044 $3.96 $0.569 pricier 1M Harder reasoning, same 1M window
GPT-5.4 mini $0.75 cache $0.075 $4.50 $0.541 pricier 400K Teams already on OpenAI tooling
Claude Sonnet 4.6 $3.00 cache $0.30 $15.00 $2.05 pricier 1M Long agentic coding sessions
Gemini 2.5 Flash $0.30 cache $0.03 $2.50 $0.272 pricier 1M Google stack, multimodal input
Grok 4.3 $1.25 cache $0.20 $2.50 $0.558 pricier 1M Grok long-context agents
§ 05 / DEEP LINKS

Specific scenarios.

All calculators →

Frequently asked.

Short answers for teams checking DeepSeek V4 Flash pricing, billing behaviour, and migration choices.

Q · 01 Why are there two prices for DeepSeek V4 Flash? +
Because DeepSeek bills by the hour of day since 16:00 UTC on August 16, 2026. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday; every other hour is off-peak at half the peak rate. For this model the peak card is $0.44/M input on a cache miss, $0.014/M on a cache hit and $1.32/M output; off-peak is $0.22 / $0.007 / $0.66. The card it replaced was flat at $0.14 / $0.0028 / $0.28, so output rose 2.4x off-peak and 4.7x at peak. Re-verified on the vendor's page on August 23, 2026: the scheme is live and the rates are unchanged since the switch.
Q · 02 How much does DeepSeek V4 Flash cost? +
DeepSeek's pricing page lists two tiers since August 16, 2026. At peak (01:00-04:00 and 06:00-10:00 UTC) it is $0.44/M cache-miss input, $0.014/M cache-hit input and $1.32/M output; off-peak, which is the other 17 hours of the day, is exactly half: $0.22 / $0.007 / $0.66. The context window is 1M tokens either way. Under AI//COST's 92/8 agentic blend with 82% cache hits, the effective planning figure is $0.048/M.
Q · 03 Which model version does deepseek-v4-flash actually serve? +
As of July 31, 2026 the id serves DeepSeek-V4-Flash-0731. DeepSeek states the calling method is unchanged — you keep requesting deepseek-v4-flash and get the newest checkpoint. We re-verified the rate card on August 2, 2026 after the swap: the prices did not move.
Q · 04 Is a price change coming? +
None is announced. Through early August the page carried an unspecified warning that DeepSeek planned to raise overall API pricing in the near future, with a significant increase expected. That increase arrived on August 16, 2026 as the peak/off-peak card above, and the warning has since been removed from the page — as of August 23, 2026 it carries only the standing note that prices may vary. So there is no scheduled change to record, and no date to watch. Worth remembering that DeepSeek's previous announced change, a May 31 reversion on V4 Pro, never happened at all: on this vendor an announcement is a signal, not a schedule.
Q · 05 Do deepseek-chat and deepseek-reasoner still work? +
No. DeepSeek retired both ids after July 24, 2026, 15:59 UTC and they are now inaccessible. Non-thinking traffic that used deepseek-chat and reasoning traffic that used deepseek-reasoner both move to deepseek-v4-flash (or deepseek-v4-pro), switching between the two behaviours via thinking mode rather than via the model id.
Q · 06 Does thinking mode cost extra? +
No. The rate card publishes one set of prices per model, and it covers both non-thinking and thinking modes — thinking is the default. Note that thinking still produces tokens, so a reasoning-heavy request bills more than a terse one at the same rate; the mode is free, the tokens are not.
Q · 07 How cheap is the cache, really? +
A cache hit is $0.014/M against $0.44/M for a miss at peak — a 31x spread, and DeepSeek states cache storage itself is free. The ratio is identical off-peak, since both halve. That is why the effective blended figure lands near $0.19/M at peak rather than near the headline input rate. If your prompts do not repeat, budget the full miss rate instead — $0.44/M at peak, $0.22/M off-peak.
Q · 08 What is the output limit? +
DeepSeek lists a maximum output of 384K tokens for both V4 Flash and V4 Pro, inside the shared 1M-token context window.
Q · 09 When should I pay for V4 Pro instead? +
V4 Pro runs about 3x the effective rate ($0.57/M against $0.19/M at peak) for the same 1M context, so the question is whether the harder reasoning earns that. The endpoint caveat that used to sit here is gone: as of August 23, 2026 DeepSeek's feature table marks the Responses API as supported on all three models, V4 Pro included, alongside the Anthropic-format endpoint. Pro's concurrency limit is the remaining gap — 500 against 2500 on Flash.