Mercury API Pricing
Inception builds diffusion LLMs — models that generate tokens in parallel rather than one after another. The Mercury family is the only architecture of its kind priced in this catalogue, and the company sells it on throughput: over 1,000 tokens per second, or as its own marketing puts it, frontier quality at five times the speed.
Mercury 2 LAUNCHED AUG 2026
Inception's generally available reasoning dLLM, billed at $0.25/M input, $0.025/M cached input and $0.75/M output on the developer pricing table.
Mercury Edit 2 LAUNCHED AUG 2026
Code-editing sibling of Mercury 2 at the identical rate card ($0.25 / $0.025 / $0.75).
Mercury 2.5 Preview LAUNCHED AUG 2026
Inception prices Mercury 2.5 Preview at $0.20/M input, $0.02/M cached input and $0.75/M output on its own models page, and describes it as "our most intelligent reasoning dLLM, at our lowest price"…
| Model | Input /M | Output /M | Cached | Context | Max output | Vision | Tools | Tier |
|---|---|---|---|---|---|---|---|---|
| Mercury 2 | $0.25 | $0.75 | $0.025−90% | 128K | — | ✗ | ✓ | Active |
| Mercury Edit 2 | $0.25 | $0.75 | $0.025−90% | 32K | — | ✗ | ✓ | Active |
| Mercury 2.5 Preview FLAGSHIP | $0.20 | $0.75 | $0.02−90% | 260K | — | ✗ | ✓ | Preview |
Pricing across the lineup.
The Mercury story so far.
Prices are read from Inception's own developer table and models page. Company history is from the launch coverage and the lead investor's announcement, cited as reported rather than vendor-stated.
platform.inceptionlabs.ai — API keys and playground. The endpoint is OpenAI-compatible, so adopting it is a base-URL and model-name change rather than a rewrite. Every new account is granted 100 million free tokens.
DOCSThe developer table that prices Mercury 2 and Mercury Edit 2, with endpoints, context windows and max output per model. This is the source for the rates on this page.
RESELLERWhere Mercury 2.5 Preview actually runs — Inception's own banner points there. Note that a reseller sets its own margin, so the rate you pay there need not equal the vendor's published figure.
Inception came out of stealth in February 2025 and is based in Palo Alto. It was founded by three academics — Stefano Ermon of Stanford, who co-invented diffusion models, Aditya Grover of UCLA and Volodymyr Kuleshov of Cornell. In November 2025 it raised a $50 million seed led by Menlo Ventures, with Mayfield, Innovation Endeavors, Microsoft's M12, Snowflake Ventures, Databricks and Nvidia's NVentures participating. Those company facts come from launch coverage and the investors' own announcements, not from a vendor pricing page.
The technical bet is what makes this provider worth a page. Every other model in this catalogue is autoregressive: it produces one token, then the next, conditioned on what came before. A diffusion LLM starts from noise and refines a whole span in parallel. The practical consequence is throughput — Inception quotes over 1,000 tokens per second for Mercury 2 — which is why the company's own framing is speed rather than price.
And that framing matters when reading the rates. At $0.20–$0.25 per million input tokens Mercury is not a budget outlier: it lands in the same band as the mainstream cheap tiers from the large labs, above the cheapest Chinese rows. Buying Mercury to cut a token bill is the wrong reason. Buying it because a task is latency-bound — voice, autocomplete, an agent loop where a slow frontier model is unusable at any price — is the argument the vendor is actually making, and the one worth testing.
One wrinkle to plan around: Mercury 2.5 Preview, the cheapest and longest-context row, is not served from Inception's own platform. The company's site banner says it is live on OpenRouter, and it is absent from the developer pricing table. The price on this site is Inception's published figure; what a reseller charges is its own business.
Google (Gemini Flash)
The mainstream fast tier most teams reach for first, with vision and a far larger ecosystem. Mercury's pitch against it is raw tokens per second, not capability breadth or price.
PRICE PEERAnthropic (Haiku)
Similar price band from a frontier lab, with vision, caching and a much wider tool ecosystem. If latency is not the binding constraint, this is the safer default.
CHEAPERAlibaba (Qwen Flash)
Materially cheaper per token and multimodal, but China-hosted, which is a procurement question Inception does not raise. Mercury does not win on price against this tier.
CODE PEERKwaipilot (KAT Coder)
Coding-specialised rows in a comparable band, aimed at the same IDE and agent workloads Mercury Edit 2 targets.
OPEN-WEIGHT GIANTAlibaba (Qwen)
The Qwen (Tongyi Qianwen) model family from Alibaba Cloud's Tongyi Lab. First released in 2023, Qwen is the most-downloaded open-weight model family in the world — most tiers ship under Apache 2.0 on Hugging Face and ModelScope, while the proprietary Max tier is served through Alibaba Cloud's Model Studio.
GLM ARCHITECTZhipu (Z.ai / GLM)
The Beijing lab behind the GLM models, also operating internationally as Z.ai. Spun out of Tsinghua University in 2019, Zhipu ships strong coding/agentic models, multiple free tiers, and open weights — and was the first of China's "Six Tigers" to pursue an IPO. It also sits on the US Entity List.
Frequently asked.
Practical questions about Mercury pricing, the free allowance, and what a diffusion LLM does and does not change about your bill.
Q · 01 Which Mercury model should I start with? +
$0.25/M input, $0.75/M output, 128K context, and 100 million free tokens on a new account. Mercury 2.5 Preview is cheaper ($0.20/M) with a 260K window, but Inception serves it through OpenRouter rather than its own platform, so it is the better model on paper and the more awkward one to integrate.Q · 02 Is Mercury actually cheap? +
$0.20–$0.25 per million input tokens it sits in the same band as the mainstream cheap tiers from the large labs, and well above the cheapest Chinese rows. What the company sells is speed — over 1,000 tokens per second — so the honest comparison is cost per completed task under a latency budget, not dollars per million tokens.Q · 03 What is a diffusion LLM? +
Q · 04 How real is the free tier? +
$12 of usage, and unlike a trial credit it does not expire into a different price the moment you finish evaluating.Q · 05 Why is a preview cheaper than the generally available model? +
$0.20 against $0.25 on input, and 260K of context against 128K. Output is identical at $0.75. A newer model being better and cheaper is ordinary; a preview being both while the GA row stays on the price list at a higher rate is less so, and it is worth watching what happens to Mercury 2's price when the preview reaches general availability.