Last verified
NO NVIDIA TOKEN PRICEFASTEST 30B TIER1M CONTEXT30B PARAMSFREE EVAL ENDPOINT

Nemotron 3.5 Lightning 30B A3B API Pricing

NVIDIA does not sell Nemotron 3.5 Lightning 30B A3B by the token. Its model page is the fastest 30B A3B MoE in the Nemotron line, aimed at specialised agentic tasks, with a playground and full specifications — but no NVIDIA price. The per-token figures shown there sit under Partner Endpoints and belong to third-party hosters: Bitdeer AI at $0.10/M input and $0.95/M output. A hoster's rate is what that hoster charges to run the weights, not what NVIDIA charges. NVIDIA's own routes are a free evaluation endpoint, a downloadable NIM microservice you self-host, and NVIDIA AI Enterprise — which publishes no figure at all and directs buyers to contact sales.

Input - per 1M tokens
Not sold per token
NVIDIA publishes no rate n/a
Output - per 1M tokens
Not sold per token
Partner rates are resale n/a
Cached input
n/a
No cache rate published
Effective - agentic blend
TBA
Needs a vendor rate
§ 03 / TOKENIZER

Paste text. See tokens.

Estimate · nemotron-tokenizer-estimate · ≈3.9 chars/token Auto-counts as you type

This is a chars-per-token approximation, not a real tokenizer. Actual tokens vary by language, code density, and tool-call overhead — counts are typically ±10–20% off for English prose, more for code or non-Latin scripts. For exact billing, use the vendor's official tokenizer.

Characters 426
Words 71
Tokens (estimated) 109 tokens
§ 04 / SHELF

Up against the shelf.

All models →
Model Input /M Output /M Effective blended Context Best for
GLM-4.7 $0.60 cache $0.11 $2.20 $0.358 priced open-weight peer 200K A model you can actually price
Solar Pro 3 $0.15 cache $0.015 $0.60 $0.0842 cheap priced tier Not published Low-cost hosted alternative
DeepSeek V4 Flash $0.44 cache $0.014 $1.32 $0.189 open weights, priced 1M Open weights with a public rate

Frequently asked.

What NVIDIA publishes for Nemotron 3.5 Lightning 30B A3B, and what it does not.

Q · 01 How much does Nemotron 3.5 Lightning 30B A3B cost per token? +
NVIDIA does not say. We checked its model page on August 16, 2026: it carries specifications, a playground and deployment options, but no NVIDIA per-token rate. This is not an oversight on our side — it is how NVIDIA sells.
Q · 02 But there are prices on the NVIDIA page — why not use those? +
Because they are not NVIDIA's. They appear under Partner Endpoints and belong to hosting companies running the weights on their own hardware: Bitdeer AI at $0.10/M input and $0.95/M output. That is a resale price, in the same category as an aggregator's.
Q · 03 So how does NVIDIA charge for it? +
Three ways, none per-token. A free evaluation endpoint on build.nvidia.com for testing. A downloadable NIM microservice you run on your own GPUs. And NVIDIA AI Enterprise, the licence that covers production NIM use — whose own product page publishes no price and points to a sales contact. If you self-host, your real cost is GPU time plus that licence, not tokens.
Q · 04 What is the model itself? +
NVIDIA describes it as the fastest 30B A3B MoE in the Nemotron line, aimed at specialised agentic tasks. Context length 1M, 30B parameters, text in and text out. Function calling is supported and so is reasoning; structured output is supported.
Q · 05 What should I price against instead? +
Any open-weight model that also has a published hosted rate — DeepSeek V4 Flash and GLM-4.7 are the closest comparisons, since both ship weights and publish a first-party price. That gives you a number to plan with; treat Nemotron's partner rates as a floor for what hosting it would cost you.