Last verified
NO NVIDIA TOKEN PRICE561B FRONTIER1M CONTEXT561B PARAMSFREE EVAL ENDPOINT

Nemotron 3 Ultra 550B A55B API Pricing

NVIDIA does not sell Nemotron 3 Ultra 550B A55B by the token. Its model page is an open hybrid Mamba-Transformer MoE for agentic reasoning, coding and planning, with a playground and full specifications — but no NVIDIA price. The per-token figures shown there sit under Partner Endpoints and belong to third-party hosters: Bitdeer AI at $0.80 / $2.60, Deep Infra at $0.50 / $2.20, and Digital Ocean from $0.90 input. Those figures disagree with each other for the same model, which is exactly why a hoster's rate is not the model's price. NVIDIA's own routes are a free evaluation endpoint, a downloadable NIM microservice you self-host, and NVIDIA AI Enterprise — which publishes no figure at all and directs buyers to contact sales.

Input - per 1M tokens
Not sold per token
NVIDIA publishes no rate n/a
Output - per 1M tokens
Not sold per token
Partner rates are resale n/a
Cached input
n/a
No cache rate published
Effective - agentic blend
TBA
Needs a vendor rate
§ 03 / TOKENIZER

Paste text. See tokens.

Estimate · nemotron-tokenizer-estimate · ≈3.9 chars/token Auto-counts as you type

This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.

Characters 422
Words 68
Tokens (estimated) 108 tokens
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
GLM-4.7 $0.60 cache $0.11 $2.20 $0.358 priced open-weight peer 200K A model you can actually price
Solar Pro 3 $0.15 cache $0.015 $0.60 $0.0842 cheap priced tier Not published Low-cost hosted alternative
DeepSeek V4 Flash $0.44 cache $0.014 $1.32 $0.189 open weights, priced 1M Open weights with a public rate

Frequently asked.

What NVIDIA publishes for Nemotron 3 Ultra 550B A55B, and what it does not.

Q · 01 How much does Nemotron 3 Ultra 550B A55B cost per token? +
NVIDIA does not say. We checked its model page on August 16, 2026: it carries specifications, a playground and deployment options, but no NVIDIA per-token rate. This is not an oversight on our side — it is how NVIDIA sells.
Q · 02 But there are prices on the NVIDIA page — why not use those? +
Because they are not NVIDIA's. They appear under Partner Endpoints and belong to hosting companies running the weights on their own hardware: Bitdeer AI at $0.80 / $2.60, Deep Infra at $0.50 / $2.20, and Digital Ocean from $0.90 input. Note that they do not agree with each other for the same model — the clearest possible demonstration that a hoster's rate reflects that hoster's margins and hardware, not the model's price.
Q · 03 So how does NVIDIA charge for it? +
Three ways, none per-token. A free evaluation endpoint on build.nvidia.com for testing. A downloadable NIM microservice you run on your own GPUs. And NVIDIA AI Enterprise, the licence that covers production NIM use — whose own product page publishes no price and points to a sales contact. If you self-host, your real cost is GPU time plus that licence, not tokens.
Q · 04 What is the model itself? +
NVIDIA describes it as an open hybrid Mamba-Transformer MoE for agentic reasoning, coding and planning. Context length 1M, 561B parameters, text in and text out. Function calling is supported and so is reasoning; structured output is not supported.
Q · 05 What should I price against instead? +
Any open-weight model that also has a published hosted rate — DeepSeek V4 Flash and GLM-4.7 are the closest comparisons, since both ship weights and publish a first-party price. That gives you a number to plan with; treat Nemotron's partner rates as a floor for what hosting it would cost you.