Provider
NVIDIA API pricing: every model compared
NVIDIA offers 5 models through its API in our data, priced from $0.0469 to $0.50 per million input tokens. The newest is Nemotron 3.5 Lightning, first listed on 2026-08-11.
- Models
- 5
- Cheapest input
- $0.0469
- Priciest input
- $0.50
- Largest context
- 262,144
- With caching
- 2 of 5
- Open weights
- all
Cheapest
Nemotron 3.5 LightningNVIDIA$0.0469 in · $0.134 outLargest context
Nemotron 3 Nano 30B A3BNVIDIA262,144 tokensNewest
Nemotron 3.5 LightningNVIDIA2026-08-11Line-up
Every NVIDIA model and price
| Model | Input / 1M | Output / 1M | Context | Released |
|---|---|---|---|---|
| Nemotron 3.5 Lightning | $0.0469 | $0.134 | 262,144 | 2026-08-11 |
| Nemotron 3 Ultra | $0.50 | $2.20 | 262,144 | 2026-06-04 |
| Nemotron 3.5 Content Safety | $0.20 | $0.20 | 131,072 | 2026-06-04 |
| Nemotron 3 Super | $0.08 | $0.45 | 262,144 | 2026-03-11 |
| Nemotron 3 Nano 30B A3B | $0.06 | $0.24 | 262,144 | 2025-12-14 |
Prices in US dollars per 1M tokens, newest first. Release dates are when our source first listed the model.
Overview
NVIDIA’s pricing at a glance
The spread between NVIDIA’s cheapest and most expensive model is large: Nemotron 3 Ultra costs 11× more per input token than Nemotron 3.5 Lightning. Picking the smallest model that does the job well is usually the biggest saving available, ahead of any discount.
2 of 5 of the line-up has a cached-input price, which helps chatbots and agents that resend the same instructions. No batch prices are published in our data. 1 of 5 accept images as input.
FAQ
NVIDIA API questions
How much does the NVIDIA API cost?
NVIDIA's 5 models range from $0.0469 to $0.50 per million input tokens, and output costs more than input on almost every model. The exact bill depends on your token volumes, which you can estimate in the LLM cost calculator.
What is NVIDIA's cheapest model?
By blended price (three parts input to one part output) it is Nemotron 3.5 Lightning, at $0.0469 input and $0.134 output per million tokens. The most expensive is Nemotron 3 Ultra.
Which NVIDIA model has the largest context window?
Nemotron 3 Nano 30B A3B, with 262,144 tokens per request.
Does NVIDIA offer prompt caching or batch discounts?
In our data, 2 of 5 of NVIDIA's models have a published cached-input price and none have a published batch price. Both can cut costs substantially for repeated prompts or work that can wait.