Last verified
NO PER-TOKEN RATENEMOTRON FAMILYOPEN WEIGHTSFREE EVAL ENDPOINTSELF-HOST NIM

NVIDIA Nemotron Pricing

NVIDIA builds the Nemotron open model family and serves it through NIM inference microservices. It is the one major lab on this site that does not sell its models by the token — you evaluate free, self-host under a licence, or buy from a partner. That makes "the price of Nemotron" a different question from the one this site usually answers.

Models tracked
3
Nemotron 3 / 3.5
NVIDIA token rate
none
not sold per token
Largest tracked
561B
Nemotron 3 Ultra
Max context
1M
Ultra + 3.5 Lightning
Eval endpoint
Free
build.nvidia.com
Production licence
Quote
AI Enterprise, contact sales
§ 01 / LINEUP

The full roster.

Side-by-side →
§ 02 / SHELF

All side-by-side.

Methodology →
Model Input /M Output /M Cached Context Max output Vision Tools Tier
Nemotron 3 Ultra 550B A55B RESTRICTED 1M Free evaluation endpoint, self-hosted NIM, or partner endpoints - NVIDIA publishes no per-token rate Restricted
Nemotron 3 Nano Omni 30B A3B Reasoning RESTRICTED 262K Free evaluation endpoint, self-hosted NIM, or partner endpoints - NVIDIA publishes no per-token rate Restricted
Nemotron 3.5 Lightning 30B A3B RESTRICTED 1M Free evaluation endpoint, self-hosted NIM, or partner endpoints - NVIDIA publishes no per-token rate Restricted

How NVIDIA sells models.

The commercial shape matters more than release dates here, because it is the reason this hub carries no token prices. Sourced from build.nvidia.com and NVIDIA's AI Enterprise product page.

AUG 16 · 2026
AI//COST adds NVIDIA — three Nemotron models recorded with pricing pending, because NVIDIA publishes no per-token rate for any of them
PRICING
AUG 16 · 2026
Checked and recorded: the per-token figures on NVIDIA's model pages belong to partner endpoints, and for Nemotron 3 Ultra three partners quote three different rates
PRICING
2026
Nemotron 3.5 generation ships, adding the Lightning 30B A3B tier for agentic work at 1M context
RELEASE
2026
Nemotron 3 Ultra 550B A55B — hybrid Mamba-Transformer MoE with 1M context, reported by NVIDIA at 52M API calls in a 30-day window
RELEASE
ONGOING
NIM microservices remain the distribution model: free evaluation endpoints, downloadable containers, and production use licensed through NVIDIA AI Enterprise
CORPORATE
§ 04 / ACCESS

Where to get it.

Methodology →
§ 04 / BEST FOR

Which NVIDIA for what.

More scenarios →
If you want to try Nemotron without paying
Free endpoint on build.nvidia.com
Open catalogue →
If you need the largest open Nemotron
Nemotron 3 Ultra 550B A55B
Profile →
If you need video, audio and image understanding
Nemotron 3 Nano Omni
Profile →
If you need a per-token price you can budget on
Pick a vendor that publishes one
Compare priced models →
§ 06 / BACKGROUND

The company behind it.

build.nvidia.com →

NVIDIA occupies an unusual position on a price-comparison site: it builds competitive open models and gives you several ways to run them, but it does not retail them by the token. That is not a gap in our data — it is the product.

The Nemotron family is distributed as NIM inference microservices. You can call a model free on build.nvidia.com to evaluate it, download the container and run it on your own GPUs, or use a partner endpoint. Production self-hosting is licensed through NVIDIA AI Enterprise, and NVIDIA's own product page for it publishes no figure at all, directing buyers to contact sales.

This creates a trap worth naming, because it catches a lot of comparison content. NVIDIA's model pages do show per-token prices — under a heading called Partner Endpoints. Those belong to hosting companies. On the Nemotron 3 Ultra page alone, Bitdeer AI quotes $0.80/$2.60, Deep Infra quotes $0.50/$2.20, and Digital Ocean starts at $0.90 input — three prices for one model, because each reflects a different company's hardware and margin. Quoting any of them as "NVIDIA's price" is the same error as quoting an aggregator.

So every Nemotron row here is recorded as pricing pending, with the specifications, the partner rates named as partner rates, and no invented figure. If you are budgeting, the honest comparison is against an open-weight model that also publishes a hosted rate — DeepSeek and Zhipu both do. If NVIDIA publishes a retail per-token rate, these rows convert to priced ones and gain calculators.

§ 07 / COMPETITORS

Other frontier labs.

All providers →

Frequently asked.

Why this hub shows no token prices, and what NVIDIA charges instead.

Q · 01 What does NVIDIA charge per million tokens for Nemotron? +
Nothing — it does not sell the models that way. We checked build.nvidia.com on August 16, 2026: the model pages carry specifications, a playground and deployment options, but no NVIDIA per-token rate. Every Nemotron row on this site is therefore marked pricing pending.
Q · 02 There are dollar figures on NVIDIA's pages. What are they? +
They are partner endpoint prices — third parties hosting the weights on their own hardware. The clearest proof is the Nemotron 3 Ultra page, where three partners quote three different rates for the same model: Bitdeer AI at $0.80/$2.60, Deep Infra at $0.50/$2.20, and Digital Ocean from $0.90 input. A number that changes with the host is a hosting price, not a model price.
Q · 03 Then how much does it actually cost to run Nemotron? +
It depends on the route. Evaluation is free on build.nvidia.com. Self-hosting costs GPU time plus an NVIDIA AI Enterprise licence, and NVIDIA publishes no price for that licence — its product page directs you to sales. Partner endpoints cost whatever that partner charges. Only the third route is per-token, and it is not NVIDIA's rate.
Q · 04 Why carry NVIDIA at all if you can't price it? +
Because "what does this cost?" is a real question with a real answer, and the answer is not per-token, here is how it is actually sold. Leaving the vendor out would send readers to aggregator numbers presented as NVIDIA's. These pages give the specifications, name the partner rates as partner rates, and refuse to invent the figure in between.
Q · 05 What should I compare against instead? +
Open-weight models that also publish a hosted price, so you get both options priced: DeepSeek V4 Flash and GLM-4.7 are the closest analogues. Treat the Nemotron partner rates as an indication of what hosting the weights costs someone else, not as a budget line.
Q · 06 Will these pages ever show prices? +
Yes, if NVIDIA publishes a retail per-token rate. These rows are recorded as pending precisely so they convert cleanly: the moment a first-party rate exists, each page gains a calculator, cost scenarios and price history like every other model here.