Nemotron 3 Ultra 550B A55B API Pricing
NVIDIA does not sell Nemotron 3 Ultra 550B A55B by the token. Its model page is an open hybrid Mamba-Transformer MoE for agentic reasoning, coding and planning, with a playground and full specifications — but no NVIDIA price. The per-token figures shown there sit under Partner Endpoints and belong to third-party hosters: Bitdeer AI at $0.80 / $2.60, Deep Infra at $0.50 / $2.20, and Digital Ocean from $0.90 input. Those figures disagree with each other for the same model, which is exactly why a hoster's rate is not the model's price. NVIDIA's own routes are a free evaluation endpoint, a downloadable NIM microservice you self-host, and NVIDIA AI Enterprise — which publishes no figure at all and directs buyers to contact sales.
Paste text. See tokens.
This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.
| Model | Input /M | Output /M | Effective blended | Context | Best for |
|---|---|---|---|---|---|
| GLM-4.7 | $0.60 cache $0.11 | $2.20 | $0.358 priced open-weight peer | 200K | A model you can actually price |
| Solar Pro 3 | $0.15 cache $0.015 | $0.60 | $0.0842 cheap priced tier | Not published | Low-cost hosted alternative |
| DeepSeek V4 Flash | $0.44 cache $0.014 | $1.32 | $0.189 open weights, priced | 1M | Open weights with a public rate |
Frequently asked.
What NVIDIA publishes for Nemotron 3 Ultra 550B A55B, and what it does not.
Q · 01 How much does Nemotron 3 Ultra 550B A55B cost per token? +
Q · 02 But there are prices on the NVIDIA page — why not use those? +
$0.80 / $2.60, Deep Infra at $0.50 / $2.20, and Digital Ocean from $0.90 input. Note that they do not agree with each other for the same model — the clearest possible demonstration that a hoster's rate reflects that hoster's margins and hardware, not the model's price.Q · 03 So how does NVIDIA charge for it? +
Q · 04 What is the model itself? +
1M, 561B parameters, text in and text out. Function calling is supported and so is reasoning; structured output is not supported.