NVIDIA Nemotron Pricing
NVIDIA builds the Nemotron open model family and serves it through NIM inference microservices. It is the one major lab on this site that does not sell its models by the token — you evaluate free, self-host under a licence, or buy from a partner. That makes "the price of Nemotron" a different question from the one this site usually answers.
Nemotron 3 Ultra 550B A55B
⚠ Free evaluation endpoint, self-hosted NIM, or partner endpoints - NVIDIA publishes no per-token rate. PRICE PENDING - NVIDIA publishes no retail per-token rate for this model.
Nemotron 3 Nano Omni 30B A3B Reasoning
⚠ Free evaluation endpoint, self-hosted NIM, or partner endpoints - NVIDIA publishes no per-token rate. PRICE PENDING - NVIDIA publishes no retail per-token rate for this model.
Nemotron 3.5 Lightning 30B A3B
⚠ Free evaluation endpoint, self-hosted NIM, or partner endpoints - NVIDIA publishes no per-token rate. PRICE PENDING - NVIDIA publishes no retail per-token rate for this model.
| Model | Input /M | Output /M | Cached | Context | Max output | Vision | Tools | Tier |
|---|---|---|---|---|---|---|---|---|
| Nemotron 3 Ultra 550B A55B RESTRICTED | — | — | — | 1M | Free evaluation endpoint, self-hosted NIM, or partner endpoints - NVIDIA publishes no per-token rate | ✗ | ✗ | Restricted |
| Nemotron 3 Nano Omni 30B A3B Reasoning RESTRICTED | — | — | — | 262K | Free evaluation endpoint, self-hosted NIM, or partner endpoints - NVIDIA publishes no per-token rate | ✓ | ✗ | Restricted |
| Nemotron 3.5 Lightning 30B A3B RESTRICTED | — | — | — | 1M | Free evaluation endpoint, self-hosted NIM, or partner endpoints - NVIDIA publishes no per-token rate | ✗ | ✗ | Restricted |
How NVIDIA sells models.
The commercial shape matters more than release dates here, because it is the reason this hub carries no token prices. Sourced from build.nvidia.com and NVIDIA's AI Enterprise product page.
NVIDIA's catalogue and playground. Generate an API key and call the model free to evaluate. This is the surface most people mean by "the NVIDIA API" — and it is not a metered commercial product.
SELF-HOST · LICENCEDownload the NIM microservice and run it on your own GPUs. Production use is licensed through NVIDIA AI Enterprise, whose product page publishes no price and directs buyers to contact sales. Your real cost here is GPU time plus that licence.
PARTNER · RESALEThird parties host the weights and charge per token — Bitdeer AI, Deep Infra, Digital Ocean, Lightning AI, Vultr. Their rates appear on NVIDIA's own model pages but are their prices, and they differ from each other for the same model.
NVIDIA occupies an unusual position on a price-comparison site: it builds competitive open models and gives you several ways to run them, but it does not retail them by the token. That is not a gap in our data — it is the product.
The Nemotron family is distributed as NIM inference microservices. You can call a model free on build.nvidia.com to evaluate it, download the container and run it on your own GPUs, or use a partner endpoint. Production self-hosting is licensed through NVIDIA AI Enterprise, and NVIDIA's own product page for it publishes no figure at all, directing buyers to contact sales.
This creates a trap worth naming, because it catches a lot of comparison content. NVIDIA's model pages do show per-token prices — under a heading called Partner Endpoints. Those belong to hosting companies. On the Nemotron 3 Ultra page alone, Bitdeer AI quotes $0.80/$2.60, Deep Infra quotes $0.50/$2.20, and Digital Ocean starts at $0.90 input — three prices for one model, because each reflects a different company's hardware and margin. Quoting any of them as "NVIDIA's price" is the same error as quoting an aggregator.
So every Nemotron row here is recorded as pricing pending, with the specifications, the partner rates named as partner rates, and no invented figure. If you are budgeting, the honest comparison is against an open-weight model that also publishes a hosted rate — DeepSeek and Zhipu both do. If NVIDIA publishes a retail per-token rate, these rows convert to priced ones and gain calculators.
DeepSeek
Open weights AND a published rate — the thing Nemotron lacks. You can self-host it or call the vendor's own API and know what a token costs.
OPEN + PRICEDZhipu (Z.ai / GLM)
Open weights, free tiers and a first-party price list, plus a far broader lineup. China-based, which is a procurement question NVIDIA does not raise.
OPEN-WEIGHT PEERMistral
EU open weights with a hosted rate card — the same self-host-or-call choice as Nemotron, but with the second option actually priced.
HOSTED FRONTIERAnthropic
No weights at all, but a transparent published rate and mature enterprise terms — the opposite trade to NVIDIA's.
OPEN-WEIGHT GIANTAlibaba (Qwen)
The Qwen (Tongyi Qianwen) model family from Alibaba Cloud's Tongyi Lab. First released in 2023, Qwen is the most-downloaded open-weight model family in the world — most tiers ship under Apache 2.0 on Hugging Face and ModelScope, while the proprietary Max tier is served through Alibaba Cloud's Model Studio.
CHINA CONSUMER LEADERByteDance (Doubao)
The Doubao model family from ByteDance — parent of TikTok and Douyin. Served through the Volcano Ark (火山方舟) platform on ByteDance's Volcano Engine cloud, Doubao powers China's most-used consumer AI app and undercuts most frontier labs on price.
Frequently asked.
Why this hub shows no token prices, and what NVIDIA charges instead.
Q · 01 What does NVIDIA charge per million tokens for Nemotron? +
build.nvidia.com on August 16, 2026: the model pages carry specifications, a playground and deployment options, but no NVIDIA per-token rate. Every Nemotron row on this site is therefore marked pricing pending.Q · 02 There are dollar figures on NVIDIA's pages. What are they? +
$0.80/$2.60, Deep Infra at $0.50/$2.20, and Digital Ocean from $0.90 input. A number that changes with the host is a hosting price, not a model price.