Kimi K3 API Pricing
Kimi K3 is Moonshot's flagship for long-horizon coding and end-to-end knowledge work, and the first Kimi model to carry a full 1M-token context window. The live Kimi pricing surface lists $3/M input and $15/M output, with cache hits at $0.30/M — roughly three times the input price of the K2 family it sits above. Thinking mode is always on, so output tokens include reasoning. Pulled directly from platform.kimi.ai daily.
Run the numbers.
Live calculator pre-loaded with current Kimi K3 rates. Tweak spend, output mix, or cache assumptions and share the URL to share the calculation.
Real-world presets.
Codebase-scale migration
Reading 100-page contracts
Ticket triage
Research planning turn
Paste text. See tokens. See cost.
This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.
| Model | Input /M | Output /M | Effective blended | Context | Best for |
|---|---|---|---|---|---|
| Kimi K3 Current | $3.00 cache $0.30 | $15.00 | $1.92 agentic 92/8 | 1M | Frontier Kimi agents - 1M context |
| Kimi K2.7 Code | $0.95 cache $0.19 | $4.00 | $0.621 cheaper | 262K | Dedicated Kimi coding work |
| Kimi K2.6 | $0.95 cache $0.16 | $4.00 | $0.598 cheaper | 262K | General multimodal Kimi production |
| Kimi K2.5 | $0.60 cache $0.10 | $3.00 | $0.415 cheaper | 262K | Cost-sensitive Kimi workloads |
| Claude Opus 4.8 | $5.00 cache $0.50 | $25.00 | $3.41 pricier | 1M | Frontier coding - long-form writing |
| GPT-5.5 | $5.00 cache $0.50 | $30.00 | $3.61 pricier | 1M | Frontier OpenAI reasoning |
| Gemini 2.5 Pro | $1.25 cache $0.125 | $10.00 | $1.10 cheaper | 2M | Long-context multimodal work |
| DeepSeek V4 Flash | $0.44 cache $0.014 | $1.32 | $0.189 cheaper | 1M | Budget reasoning and coding |
Frequently asked.
Practical pricing questions, separated from calculator assumptions and regional taxes.
Q · 01 What is Kimi K3 priced at? +
$3.00/M cache-miss input, $0.30/M cache-hit input and $15.00/M output, for a 1,048,576-token context window. Under AI//COST's 92/8 agentic blend with 82% cache hits, the effective planning figure is $1.92/M. Prices exclude local taxes, which the vendor calculates at checkout.Q · 02 Has Moonshot officially announced Kimi K3? +
platform.kimi.ai with no launch post behind them, so we published them as vendor-listed but pre-announcement. Moonshot has since published the launch post on kimi.com/blog, put a launch banner across the API docs, and made K3 the flagship on moonshot.ai, which describes it as a 2.8T-parameter natively multimodal model. Re-verified on July 28, 2026: the prices did not change with the announcement.Q · 03 Why is Kimi K3 so much more expensive than Kimi K2.6? +
3× more on input ($3/M vs $0.95/M) and 3.75× more on output ($15/M vs $4/M). On our agentic blend that is $1.92/M effective against $0.60/M for K2.6. If your workload does not need the 1M context window or K3's reasoning depth, Kimi K2.6 remains substantially cheaper.Q · 04 How does prompt caching work? +
$0.30/M instead of the fresh-input $3/M rate, a 90% discount, and the calculator's default agentic blend assumes an 82% cache-hit rate.Q · 05 How is the effective price calculated? +
$0.786/M and an effective blended cost of $1.92/M. The same formula is applied to every model in the shelf above, so the numbers are directly comparable.Q · 06 Does thinking mode change what I pay? +
reasoning_effort: max, so reasoning tokens are billed as output at $15/M. There is no cheaper non-thinking variant of K3, unlike the K2 family. Output-heavy workloads will feel this more than input-heavy ones.Q · 07 Does this include image and video input? +
Q · 08 How accurate is the tokenizer estimate? +
moonshot-tokenizer-estimate chars-per-token estimate for English text. Actual billing comes from Kimi API usage fields and can differ for Chinese, code, images, video, or mixed-language prompts.