DeepSeek API Pricing
The Chinese lab behind the DeepSeek open-weight models. Spun out of quant hedge fund High-Flyer in 2023 by Liang Wenfeng, it shocked the industry by matching frontier reasoning quality at a fraction of the cost — and shipping the weights under an MIT license.
DeepSeek V4 Flash
DeepSeek V4 Flash moved to peak/off-peak billing at 16:00 UTC on 2026-08-16, announced on DeepSeek's pricing page and applied here on that date.
DeepSeek V4 Pro
DeepSeek V4 Pro moved to peak/off-peak billing at 16:00 UTC on 2026-08-16, announced on DeepSeek's pricing page and applied here on that date.
DeepSeek V4 Flash Vision (Exp)
DeepSeek's first vision model, priced on exactly the same card as deepseek-v4-flash: peak $0.44/M input on a cache miss, $0.014/M on a cache hit, $1.32/M output; off-peak is half of each ($0.22/$0.…
DeepSeek V3 LAUNCHED DEC 2024
Original DeepSeek V3 (model name deepseek-chat at launch).
DeepSeek R1 LAUNCHED JAN 2025
First-generation reasoning model (deepseek-reasoner at launch).
| Model | Input /M | Output /M | Cached | Context | Max output | Vision | Tools | Tier |
|---|---|---|---|---|---|---|---|---|
| DeepSeek V4 Flash | $0.44 | $1.32 | $0.014−97% | 1M | — | ✗ | ✓ | Active |
| DeepSeek V4 Pro FLAGSHIP | $1.32 | $3.96 | $0.044−97% | 1M | — | ✗ | ✓ | Frontier |
| DeepSeek V4 Flash Vision (Exp) PREVIEW | $0.44 | $1.32 | $0.014−97% | 1M | — | ✓ | ✓ | Preview |
| DeepSeek V3 RETIRED | — | — | — | Retired Apr 26, 2026 | → deepseek-v4-flash | ✗ | ✗ | Retired |
| DeepSeek R1 RETIRED | — | — | — | Retired Apr 26, 2026 | → deepseek-v4-pro | ✗ | ✗ | Retired |
Pricing across the lineup.
The DeepSeek story so far.
A short, loud history: V3 → R1 → V4, with V4 arriving 484 days after V3. Sourced from DeepSeek's API docs, model cards, and contemporaneous coverage.
deepseek-v4-flash begins serving DeepSeek-V4-Flash-0731; same API id, same pricesdeepseek-chat and deepseek-reasoner become inaccessible at 15:59 UTC; V4 Flash and V4 Pro are the only chat ids leftdeepseek-chat/deepseek-reasoner fold into V4; V3 + R1 retiredplatform.deepseek.com — direct API with an OpenAI-compatible endpoint, dev console, and billing. Prices are list-low; cache-hit input is ~1/10 of base.
OPEN-WEIGHT · SELF-HOSTV4 Pro + V4 Flash weights ship under the MIT license on Hugging Face — fully self-hostable with no per-token fee and no China-routing of your data.
CONSUMER · APPchat.deepseek.com plus iOS/Android apps — free consumer chat that briefly topped the US App Store in early 2025. No API; for end-user usage.
THIRD-PARTY · HOSTSBecause the weights are MIT-licensed, third-party inference hosts (Together, Fireworks, and others) serve DeepSeek V4 from US/EU regions — useful when China data residency is a blocker.
DeepSeek was founded in July 2023 in Hangzhou by Liang Wenfeng, who also co-founded and runs the quantitative hedge fund High-Flyer (started in 2015–2016). DeepSeek is owned and funded by High-Flyer rather than by external venture capital — an unusual structure that let it train large models on a GPU cluster the fund had already built for trading.
The lab's signature is compute efficiency. Its mixture-of-experts (MoE) training recipes delivered frontier-class quality at a small fraction of the budgets reported by US labs. Liang has reportedly held a controlling personal stake (~84% as of 2024), and the team is famously lean — on the order of ~150 people, with many hired straight out of university.
DeepSeek's breakout moment came with V3 (December 2024) and then R1 (January 2025): R1 matched OpenAI o1-class reasoning at roughly 1/27 the price, the consumer app briefly topped the US App Store, and the release triggered a sharp sell-off in AI hardware stocks. In April 2026 the lab shipped the V4 family — a 1.6T-parameter MoE (V4 Pro) and a 284B MoE (V4 Flash), both with 1M-token context and MIT-licensed open weights.
Pricing is the headline, and since August 16, 2026 it depends on the hour: V4 Pro runs at $1.32/$3.96 per M at peak and $0.66/$1.98 off-peak. The old flat $0.435/$0.87 was a 75% launch promo that was supposed to revert to $1.74/$3.48 on May 31, 2026 but never did — and V4 Flash at $0.44/$1.32 peak, $0.22/$0.66 off-peak — still well below frontier US models, with cache-hit input at about 1/31 of a miss.
The trade-offs are real: the lab was text-only until August 2026 and is barely past it — the one model that reads images, V4 Flash Vision, ships under an experimental id, and nothing here handles audio. The company is China-based with data stored in China by default, and US export-control and procurement-policy questions apply. The MIT weights are the escape hatch — teams that need US/EU residency can self-host. Versus OpenAI and Anthropic, DeepSeek trades multimodal breadth and Western data governance for radically lower cost and full open weights.
Alibaba (Qwen)
Broad open-weight Qwen lineup with strong multilingual + multimodal coverage and Alibaba Cloud distribution. Comparable Chinese-data-residency questions apply.
PRIMARY RIVALOpenAI
Frontier benchmarks + multimodal + ecosystem (GPT-5, ChatGPT, Azure). But ~10× pricier input, closed weights, and US data residency.
FRONTIER LEADERAnthropic
Leads coding + reasoning with Claude and zero-retention defaults. But far higher prices, no open weights, and no China-cost story.
OPEN-WEIGHT PEERMistral
EU-native open weights (Apache 2.0) with GDPR-first data residency. Comparable openness, but trails DeepSeek on raw cost-per-token.
Meta
Llama open weights at huge scale and a Western governance posture. But no managed low-cost API like DeepSeek's, and a different licensing model.
Zhipu (Z.ai / GLM)
The Beijing lab behind the GLM models, also operating internationally as Z.ai. Spun out of Tsinghua University in 2019, Zhipu ships strong coding/agentic models, multiple free tiers, and open weights — and was the first of China's "Six Tigers" to pursue an IPO. It also sits on the US Entity List.
Frequently asked.
Practical questions about DeepSeek pricing, open weights, and data residency.
Q · 01 Which DeepSeek model should I start with? +
$0.44/$1.32 peak, $0.22/$0.66 off-peak) — it covers both non-thinking and thinking modes cheaply. Step up to V4 Pro ($1.32/$3.96 peak, $0.66/$1.98 off-peak) for the hardest reasoning. Both ship a 1M-token context. See the use case picker above.Q · 02 How much cheaper is DeepSeek than OpenAI or Anthropic? +
$1.32 vs $5) and ~7.6× cheaper on output ($3.96 vs $30). Off-peak, which is 17 hours of every weekday and all weekend, both gaps double. With cache-hit input at $0.044/M the gap widens further still. The honest caveat used to be that DeepSeek is text-only; since August 2026 there is one exception, the experimental V4 Flash Vision, and it is still narrower than what Google or OpenAI accept.Q · 03 Is there an off-peak discount? +
01:00-04:00 and 06:00-10:00 UTC, Monday to Friday, and exactly half that in every other hour, weekends included. So off-peak covers 17 hours of a weekday and all of Saturday and Sunday. We store the peak figures because DeepSeek describes off-peak as "half the peak rates" — the peak card is the list price. Anything that can be scheduled overnight or at the weekend costs half. Re-verified on the vendor page August 23, 2026.Q · 04 Are DeepSeek models open-weight? +
Q · 05 Where is my data stored, and is that a problem? +
Q · 06 Did the 75% V4 Pro promo ever end? +
$1.74/M input and $3.48/M output. That date passed without a reprice: we have re-verified the pricing page repeatedly since. On August 16, 2026 DeepSeek replaced the flat card with peak/off-peak billing at $1.32/$3.96 and $0.66/$1.98. We treat it as the standard rate rather than a promo, and we track the page daily in case it moves.