Last verified
DIFFUSION LLM1000+ TOKENS/SEC128K CONTEXT100M FREE TOKENSTOOL CALLING

Mercury 2 API Pricing

Mercury 2 is Inception's generally available diffusion LLM, billed at $0.25/M input, $0.025/M cached input and $0.75/M output with a 128K context window and 50,000 max output tokens. The vendor calls it its "first enterprise ready reasoning dLLM with over 1000 tokens / sec" - the product is speed, and the rate card sits in the same band as the large labs' cheap tiers rather than below them.

Input - per 1M tokens
$0.25/M
Platform inceptionlabs.ai GA row
Output - per 1M tokens
$0.75/M
Same across Mercury flat
Cached input - per 1M tokens
$0.025/M
10% of fresh input -90%
Effective - agentic blend
$0.12/M
92/8 split - 82% cache
§ 01 / TERMINAL

Run the numbers.

Live calculator on the developer pricing table's rate. New accounts start with 100 million free tokens, so the first sizeable workload here may cost nothing at all - model the steady state, not the trial.

$ /mo
Workload split
Prompt cache hit rate
Tokens you can process
Words equivalent (English)
Effective rate
Open full calculator (all models · share URL · CSV) →
§ 02 / SCENARIOS

Real-world presets.

§ 03 / TOKENIZER

Paste text. See tokens. See cost.

Estimate · inception-tokenizer-estimate · ≈3.85 chars/token Auto-counts as you type

This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.

Characters 539
Words 83
Tokens (estimated) 140 tokens
Cost as input · uncached $0.00003 USD
Cost as output · uncached $0.00011 USD
Cost as cached input $0.000003 USD
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
Mercury 2 Current $0.25 cache $0.025 $0.75 $0.12 agentic 92/8 128K Latency-sensitive production traffic
Mercury 2.5 Preview $0.20 cache $0.02 $0.75 $0.108 cheaper, longer, preview only 260K Reasoning at the lowest Mercury rate
Mercury Edit 2 $0.25 cache $0.025 $0.75 $0.12 same card, editing endpoints 32K Autocomplete and next-edit
Gemini 3.7 Flash $0.75 cache $0.075 $3.75 $0.481 cross-vendor speed tier 1M Google's fast mainstream model
Claude Haiku 4.5 $1.00 cache $0.10 $5.00 $0.682 cross-vendor cheap tier 200K High-volume support and classification

Frequently asked.

The rate card, the free allowance, and what you are actually paying for with a diffusion model.

Q · 01 What does Mercury 2 cost? +
Inception's developer pricing table lists $0.25/M input, $0.025/M cached input and $0.75/M output. Under this site's 92/8 agentic blend at an 82% cache-hit rate that is $0.12/M effective. The company's marketing models page states the same three figures, which is the cross-check we run on any new vendor.
Q · 02 What is the free allowance? +
Every new Inception Platform account includes 100 million free tokens. Denominating a free tier in tokens rather than dollars is rare and makes it directly comparable: at Mercury 2's blended rate, 100M tokens is roughly $12 of usage.
Q · 03 How does it compare with Mercury 2.5 Preview? +
Mercury 2.5 Preview is cheaper on input ($0.20 against $0.25) and carries a 260K window against this model's 128K, but it is a preview and Inception serves it through OpenRouter rather than its own platform. Mercury 2 is the row you can call on the vendor's API today.
Q · 04 What is Mercury Edit 2 for? +
Mercury Edit 2 carries the identical rate card but is not a chat model: it serves fill-in-the-middle and next-edit endpoints for code tooling, with a 32K window. Inception notes it is supported for existing customers rather than pitched to new ones.
Q · 05 Does the speed claim change the economics? +
Only through what it replaces. Billing is per token as usual, so a faster model does not lower the token bill. What it changes is the cost of latency - interactive work, voice, autocomplete - where a slow frontier model is unusable at any price. Compare on tokens against Haiku 4.5 or Gemini 3.7 Flash, then decide whether the throughput is worth it.