DeepSeek V4 Pro API Pricing
DeepSeek bills by time of day from 16:00 UTC on August 16, 2026. Peak hours are 01:00-04:00 and 06:00-10:00 UTC; every other hour is off-peak at exactly half the peak rate. The tiles below carry the peak rate, because DeepSeek describes off-peak as "half the peak rates" - peak is the list price and off-peak is the discount. DeepSeek V4 Pro is the DeepSeek thinking flagship, and at peak it is $1.32/M input, $3.96/M output and $0.044/M cached input; off-peak it is $0.66 / $1.98 / $0.022. That replaced a flat $0.435 / $0.87 card which had itself survived an announced reversion to $1.74/$3.48 that DeepSeek never applied - so this is the first rise that actually landed.
Run the numbers.
Calculator pre-loaded with DeepSeek V4 Pro PEAK rates — halve every result for an off-peak run, which covers 17 of every 24 hours. Tweak spend, output mix, or cache hit rate to compare this model with nearby alternatives.
Real-world presets.
Repo-wide bug fix
Reading 100-page contracts
Support agent ticket triage
Research planning turn
Paste text. See tokens. See cost.
Counts use a chars-per-token calibration measured on the vendor's own published tokenizer (deepseek-ai/DeepSeek-V3.2-Exp, 2026-06-10). English prose is typically within a few percent; code and non-Latin scripts tokenize heavier. For billing-exact counts use the vendor's count-tokens API.
| Model | Input /M | Output /M | Effective blended | Context | Best for |
|---|---|---|---|---|---|
| DeepSeek V4 Pro | $1.32 cache $0.044 | $3.96 | $0.569 current page - peak | 1M | Low-cost reasoning |
| DeepSeek V4 Flash | $0.44 cache $0.014 | $1.32 | $0.189 cheaper | 1M | Cheap DeepSeek workloads |
| GPT-5.4 mini | $0.75 cache $0.075 | $4.50 | $0.541 pricier | 400K | Open-weight multimodal work |
| Claude Sonnet 4.6 | $3.00 cache $0.30 | $15.00 | $2.05 pricier | 1M | Open-weight multimodal work |
| Gemini 2.5 Flash | $0.30 cache $0.03 | $2.50 | $0.272 pricier | 1M | Open-weight multimodal work |
| Grok 4.3 | $1.25 cache $0.20 | $2.50 | $0.558 pricier | 1M | Grok long-context agents |
By comparison
Frequently asked.
Short answers for teams checking DeepSeek V4 Pro pricing, billing behaviour, and how it differs from V4 Flash.
Q · 01 How much does DeepSeek V4 Pro cost? +
Q · 02 Is this a promotional price that will expire? +
$1.74/$3.48 after May 31, 2026. That date passed with no reprice; we re-verified the pricing page on August 16, 2026, when peak/off-peak billing replaced that card with $1.32/$3.96 and $0.66/$1.98 — this is what the vendor charges. We now treat it as the standard list rate rather than a countdown, and we re-check daily.Q · 03 Is a price change coming? +
Q · 04 How does it differ from V4 Flash? +
$0.57/M against $0.19/M at peak — for the same 1M context and the same 384K max output. Both support thinking and non-thinking modes at one price. Buy Pro for the harder reasoning, not for the context window.Q · 05 Does thinking mode cost extra? +
Q · 06 Can I call it through the OpenAI Responses API? +
Q · 07 Do deepseek-chat and deepseek-reasoner still work? +
deepseek-reasoner now goes to deepseek-v4-pro or deepseek-v4-flash with thinking mode on.