Last verified
DIFFUSION LLM260K CONTEXTPREVIEW, NOT GACHEAPEST MERCURYOPENROUTER ONLY

Mercury 2.5 Preview API Pricing

Mercury 2.5 Preview is a diffusion LLM - it generates tokens in parallel rather than one after another, which is what lets Inception quote speeds above 1,000 tokens per second. It is listed at $0.20/M input, $0.02/M cached input and $0.75/M output, making it both cheaper and longer than the generally available Mercury 2. The catch is where you call it: Inception's own banner says it is live on OpenRouter, and it does not appear on the company's developer pricing table at all.

Input - per 1M tokens
$0.20/M
-20% on input vs Mercury 2
Output - per 1M tokens
$0.75/M
Same as Mercury 2 unchanged
Cached input - per 1M tokens
$0.02/M
10% of fresh input -90%
Effective - agentic blend
$0.108/M
92/8 split - 82% cache
§ 01 / TERMINAL

Run the numbers.

Live calculator on Inception's published rate. Worth noting before you model anything: the number this vendor competes on is latency, not price - the token rates here sit close to the cheap tier of every major lab, and the claim being made is throughput.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Estimate · inception-tokenizer-estimate · ≈3.85 chars/token Auto-counts as you type

This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.

Characters 608
Words 97
Tokens (estimated) 158 tokens
Cost as input · uncached $0.00003 USD
Cost as output · uncached $0.00012 USD
Cost as cached input $0.000003 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Mercury 2.5 Preview Current $0.20 cache $0.02 $0.75 $0.108 agentic 92/8 260K Reasoning and long-context work at the lowest Mercury rate
Mercury 2 $0.25 cache $0.025 $0.75 $0.12 the GA row it undercuts 128K Production traffic on Inception's own platform
Mercury Edit 2 $0.25 cache $0.025 $0.75 $0.12 same card, editing endpoints 32K Autocomplete and next-edit
Gemini 3.7 Flash $0.75 cache $0.075 $3.75 $0.481 cross-vendor speed tier 1M Google's fast mainstream model
Claude Haiku 4.5 $1.00 cache $0.10 $5.00 $0.682 cross-vendor cheap tier 200K High-volume support and classification

Frequently asked.

Where you can actually call it, why a preview costs less than the GA model, and what a diffusion LLM changes about the bill.

Q · 01 What does Mercury 2.5 Preview cost? +
Inception lists $0.20/M input, $0.02/M cached input and $0.75/M output on its models page. Cached input is 10% of fresh input, the same ratio as the rest of the Mercury line. Under this site's 92/8 agentic blend at an 82% cache-hit rate that works out to $0.108/M effective.
Q · 02 Where do I call it? It is not on the developer pricing page. +
Correct, and that is the oddity worth knowing before you plan around it. Inception's documentation prices only Mercury 2 and Mercury Edit 2; a banner across the vendor's site reads "Mercury 2.5 (Preview) is live on OpenRouter". So the model is reached through a reseller while the price comes from the vendor. The figure stored here is Inception's own published rate - we never copy a reseller's resale price.
Q · 03 Why is the preview cheaper than the generally available model? +
Inception describes it as "our most intelligent reasoning dLLM, at our lowest price". Input is $0.20 against Mercury 2's $0.25 and the context window is 260K against 128K, while output is identical at $0.75. On our blend that is about 10% cheaper overall. A newer model that is better and cheaper is normal; a preview that is both, while the GA row stays on the price list, is less usual.
Q · 04 What is a diffusion LLM, and does it change the cost model? +
It generates tokens in parallel rather than strictly left to right. It does not change how you are billed - still per input and output token - but it changes what the vendor is selling. Inception quotes over 1,000 tokens per second and markets "frontier LLM quality at 5x greater speed", so the case for this row is latency per task, not dollars per million tokens, where it sits alongside the cheap tiers of the large labs rather than beneath them.
Q · 05 Is there a free allowance? +
Yes, on Inception's own platform: every new account is granted 100 million free tokens. That is an allowance in tokens rather than a dollar credit, which makes it unusually easy to size - but it applies to the platform, and this model is served through OpenRouter.