Last verified
DIFFUSION LLMs1000+ TOKENS/SECPALO ALTO100M FREE TOKENSOPENAI-COMPATIBLE API

Mercury API Pricing

Inception builds diffusion LLMs — models that generate tokens in parallel rather than one after another. The Mercury family is the only architecture of its kind priced in this catalogue, and the company sells it on throughput: over 1,000 tokens per second, or as its own marketing puts it, frontier quality at five times the speed.

Models priced here
3
2 GA · 1 preview
Founded
2024
Out of stealth Feb 2025
Cheapest tier
$0.2/M
Mercury 2.5 Preview input
Free allowance
100M
Tokens per new account
Context window
260K
Mercury 2.5 Preview
Seed funding
$50M
Led by Menlo Ventures
§ 01 / LINEUP

The full roster.

Side-by-side →
§ 02 / SHELF

All side-by-side.

Methodology →
Model Input /M Output /M Cached Context Max output Vision Tools Tier
Mercury 2 $0.25 $0.75 $0.025−90% 128K Active
Mercury Edit 2 $0.25 $0.75 $0.025−90% 32K Active
Mercury 2.5 Preview FLAGSHIP $0.20 $0.75 $0.02−90% 260K Preview
§ 03 / PRICE CURVE

Pricing across the lineup.

How Inception (Mercury) priced 3 models · AUG 26.

oldest → newest →
$1$0.75$0.5$0.25 $0.75 $0.2 $0.75 $0.25 $0.75 $0.25 Mercury 2.5 Pre… AUG 26 Mercury 2 AUG 26 Mercury Edit 2 AUG 26
Input · newest $0.25/M
Output · newest $0.75/M
Each point is a model at its listed $/M price.

The Mercury story so far.

Prices are read from Inception's own developer table and models page. Company history is from the launch coverage and the lead investor's announcement, cited as reported rather than vendor-stated.

SEP 02 · 2026
AI//COST adds Inception — three Mercury rows priced from the vendor's own table, after the model radar flagged a provider we did not carry
PRICING
AUG 31 · 2026
Mercury 2.5 Preview appears — priced by Inception at $0.20/M, below the GA model, but served through OpenRouter rather than Inception's platform
RELEASE
NOV · 2025
$50M seed led by Menlo Ventures, with Mayfield, Innovation Endeavors, M12, Snowflake Ventures, Databricks and NVentures participating
FUNDING
FEB 26 · 2025
Inception emerges from stealth with Mercury, presented as the first commercial-scale diffusion LLM
RELEASE
§ 04 / ACCESS

Where to get it.

Methodology →
§ 04 / BEST FOR

Which Inception (Mercury) for what.

More scenarios →
If you want the cheapest and longest Mercury
Mercury 2.5 Preview
Profile →
If you need it on Inception's own API
Mercury 2
Profile →
If you are wiring IDE autocomplete
Mercury Edit 2
Profile →
If price per token is the only axis…
Compare the cheap tiers first
All models →
§ 06 / BACKGROUND

The company behind it.

www.inceptionlabs.ai →

Inception came out of stealth in February 2025 and is based in Palo Alto. It was founded by three academics — Stefano Ermon of Stanford, who co-invented diffusion models, Aditya Grover of UCLA and Volodymyr Kuleshov of Cornell. In November 2025 it raised a $50 million seed led by Menlo Ventures, with Mayfield, Innovation Endeavors, Microsoft's M12, Snowflake Ventures, Databricks and Nvidia's NVentures participating. Those company facts come from launch coverage and the investors' own announcements, not from a vendor pricing page.

The technical bet is what makes this provider worth a page. Every other model in this catalogue is autoregressive: it produces one token, then the next, conditioned on what came before. A diffusion LLM starts from noise and refines a whole span in parallel. The practical consequence is throughput — Inception quotes over 1,000 tokens per second for Mercury 2 — which is why the company's own framing is speed rather than price.

And that framing matters when reading the rates. At $0.20–$0.25 per million input tokens Mercury is not a budget outlier: it lands in the same band as the mainstream cheap tiers from the large labs, above the cheapest Chinese rows. Buying Mercury to cut a token bill is the wrong reason. Buying it because a task is latency-bound — voice, autocomplete, an agent loop where a slow frontier model is unusable at any price — is the argument the vendor is actually making, and the one worth testing.

One wrinkle to plan around: Mercury 2.5 Preview, the cheapest and longest-context row, is not served from Inception's own platform. The company's site banner says it is live on OpenRouter, and it is absent from the developer pricing table. The price on this site is Inception's published figure; what a reseller charges is its own business.

§ 07 / COMPETITORS

Other frontier labs.

All providers →
SPEED PEER

Google (Gemini Flash)

The mainstream fast tier most teams reach for first, with vision and a far larger ecosystem. Mercury's pitch against it is raw tokens per second, not capability breadth or price.

Gemini 3.7 Flash
PRICE PEER

Anthropic (Haiku)

Similar price band from a frontier lab, with vision, caching and a much wider tool ecosystem. If latency is not the binding constraint, this is the safer default.

Haiku 4.5 $1/$5
CHEAPER

Alibaba (Qwen Flash)

Materially cheaper per token and multimodal, but China-hosted, which is a procurement question Inception does not raise. Mercury does not win on price against this tier.

Qwen3.8-Flash
CODE PEER

Kwaipilot (KAT Coder)

Coding-specialised rows in a comparable band, aimed at the same IDE and agent workloads Mercury Edit 2 targets.

KAT Coder Air
OPEN-WEIGHT GIANT

Alibaba (Qwen)

The Qwen (Tongyi Qianwen) model family from Alibaba Cloud's Tongyi Lab. First released in 2023, Qwen is the most-downloaded open-weight model family in the world — most tiers ship under Apache 2.0 on Hugging Face and ModelScope, while the proprietary Max tier is served through Alibaba Cloud's Model Studio.

33 models · from $0.03/M
GLM ARCHITECT

Zhipu (Z.ai / GLM)

The Beijing lab behind the GLM models, also operating internationally as Z.ai. Spun out of Tsinghua University in 2019, Zhipu ships strong coding/agentic models, multiple free tiers, and open weights — and was the first of China's "Six Tigers" to pursue an IPO. It also sits on the US Entity List.

16 models · from $0.07/M

Frequently asked.

Practical questions about Mercury pricing, the free allowance, and what a diffusion LLM does and does not change about your bill.

Q · 01 Which Mercury model should I start with? +
Mercury 2 if you want Inception's own API — $0.25/M input, $0.75/M output, 128K context, and 100 million free tokens on a new account. Mercury 2.5 Preview is cheaper ($0.20/M) with a 260K window, but Inception serves it through OpenRouter rather than its own platform, so it is the better model on paper and the more awkward one to integrate.
Q · 02 Is Mercury actually cheap? +
Not especially, and Inception does not claim it is. At $0.20–$0.25 per million input tokens it sits in the same band as the mainstream cheap tiers from the large labs, and well above the cheapest Chinese rows. What the company sells is speed — over 1,000 tokens per second — so the honest comparison is cost per completed task under a latency budget, not dollars per million tokens.
Q · 03 What is a diffusion LLM? +
A model that refines a whole span of text in parallel rather than emitting one token at a time. Every other model priced on this site is autoregressive. Billing is unchanged — you still pay per input and output token — but the throughput profile is different, which is the entire reason to consider it.
Q · 04 How real is the free tier? +
Unusually concrete: 100 million tokens per new account, denominated in tokens rather than dollars. At Mercury 2's blended rate that is roughly $12 of usage, and unlike a trial credit it does not expire into a different price the moment you finish evaluating.
Q · 05 Why is a preview cheaper than the generally available model? +
Inception describes Mercury 2.5 Preview as "our most intelligent reasoning dLLM, at our lowest price" — $0.20 against $0.25 on input, and 260K of context against 128K. Output is identical at $0.75. A newer model being better and cheaper is ordinary; a preview being both while the GA row stays on the price list at a higher rate is less so, and it is worth watching what happens to Mercury 2's price when the preview reaches general availability.
Q · 06 Is Mercury open weight? +
No. Inception publishes no downloadable weights; access is through its API or through resellers. That is a difference from the open-weight Chinese labs in the same price band, where self-hosting is a fallback if the vendor's terms or hosting location become a problem.