Last verified
VERIFIED PRICETHINKING FLAGSHIP1MTEXT MODELPROMPT CACHING

DeepSeek V4 Pro API Pricing

DeepSeek bills by time of day from 16:00 UTC on August 16, 2026. Peak hours are 01:00-04:00 and 06:00-10:00 UTC; every other hour is off-peak at exactly half the peak rate. The tiles below carry the peak rate, because DeepSeek describes off-peak as "half the peak rates" - peak is the list price and off-peak is the discount. DeepSeek V4 Pro is the DeepSeek thinking flagship, and at peak it is $1.32/M input, $3.96/M output and $0.044/M cached input; off-peak it is $0.66 / $1.98 / $0.022. That replaced a flat $0.435 / $0.87 card which had itself survived an announced reversion to $1.74/$3.48 that DeepSeek never applied - so this is the first rise that actually landed.

Input - per 1M tokens
$1.32/M
Off-peak $0.66/M peak
Output - per 1M tokens
$3.96/M
Off-peak $1.98/M peak
Cached input
$0.044/M
Off-peak $0.022/M peak
Effective - agentic blend
$0.569/M
92/8 split - 82% cache
§ 01 / TERMINAL

Run the numbers.

Calculator pre-loaded with DeepSeek V4 Pro PEAK rates — halve every result for an off-peak run, which covers 17 of every 24 hours. Tweak spend, output mix, or cache hit rate to compare this model with nearby alternatives.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Calibrated · measured on the vendor's tokenizer · 2026-06-10 Auto-counts as you type

Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (deepseek-ai/DeepSeek-V3.2-Exp, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.

Characters 344
Words 56
Tokens (estimated) 65 tokens
Cost as input · uncached $0.00009 USD
Cost as output · uncached $0.00026 USD
Cost as cached input $0.000003 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
DeepSeek V4 Pro $1.32 cache $0.044 $3.96 $0.569 current page - peak 1M Low-cost reasoning
DeepSeek V4 Flash $0.44 cache $0.014 $1.32 $0.189 cheaper 1M Cheap DeepSeek workloads
GPT-5.4 mini $0.75 cache $0.075 $4.50 $0.541 pricier 400K Open-weight multimodal work
Claude Sonnet 4.6 $3.00 cache $0.30 $15.00 $2.05 pricier 1M Open-weight multimodal work
Gemini 2.5 Flash $0.30 cache $0.03 $2.50 $0.272 pricier 1M Open-weight multimodal work
Grok 4.3 $1.25 cache $0.20 $2.50 $0.558 pricier 1M Grok long-context agents
§ 05 / DEEP LINKS

Specific scenarios.

All calculators →

Frequently asked.

Short answers for teams checking DeepSeek V4 Pro pricing, billing behaviour, and how it differs from V4 Flash.

Q · 01 How much does DeepSeek V4 Pro cost? +
DeepSeek lists two tiers since August 16, 2026. At peak (01:00-04:00 and 06:00-10:00 UTC) it is $1.32/M cache-miss input, $0.044/M cache-hit input and $3.96/M output; off-peak — the other 17 hours — is exactly half: $0.66 / $0.022 / $1.98. The context window is 1M tokens either way. Under AI//COST's 92/8 agentic blend with 82% cache hits, the effective planning figure is $0.144/M.
Q · 02 Is this a promotional price that will expire? +
It arrived as one, and the expiry did not happen. DeepSeek launched V4 Pro at a 75% discount and said the price would revert to $1.74/$3.48 after May 31, 2026. That date passed with no reprice; we re-verified the pricing page on August 16, 2026, when peak/off-peak billing replaced that card with $1.32/$3.96 and $0.66/$1.98 — this is what the vendor charges. We now treat it as the standard list rate rather than a countdown, and we re-check daily.
Q · 03 Is a price change coming? +
None is announced. Through early August the page carried an unspecified warning that DeepSeek planned to raise overall API pricing in the near future, with a significant increase expected. That increase landed on August 16, 2026 as the peak/off-peak card above, and the warning has since been removed — as of August 23, 2026 the page carries only the standing note that prices may vary. Nothing is scheduled, so there is no date to watch. Keep in mind that this vendor's previous announced change, the May 31 reversion on this very model, never happened: on DeepSeek an announcement is a signal, not a schedule.
Q · 04 How does it differ from V4 Flash? +
V4 Pro is the 1.6T-parameter MoE (49B active) against V4 Flash's 284B (13B active), and it runs about the effective rate — $0.57/M against $0.19/M at peak — for the same 1M context and the same 384K max output. Both support thinking and non-thinking modes at one price. Buy Pro for the harder reasoning, not for the context window.
Q · 05 Does thinking mode cost extra? +
No. One rate card covers both non-thinking and thinking modes, and thinking is the default. The mode is free; the tokens it generates are billed at the normal output rate, so a reasoning-heavy request still costs more than a terse one.
Q · 06 Can I call it through the OpenAI Responses API? +
Yes, as of August 23, 2026. DeepSeek's feature table now marks the Responses API as supported on all three models, V4 Pro included — the gap that existed when this page was first written has closed. The OpenAI-compatible Chat Completions endpoint and the Anthropic-format endpoint work with V4 Pro too. The remaining difference from V4 Flash is throughput, not protocol: Pro is limited to 500 concurrent requests against Flash's 2500.
Q · 07 Do deepseek-chat and deepseek-reasoner still work? +
No. Both ids were fully retired after July 24, 2026, 15:59 UTC. Reasoning traffic that used deepseek-reasoner now goes to deepseek-v4-pro or deepseek-v4-flash with thinking mode on.