API pricing

# NVIDIA API pricing

> NVIDIA API prices run from $0.05 per million input tokens (Nemotron 3 Nano 30B A3B) to $0.50 (Nemotron 3 Ultra). Its highest-ranked model, Nemotron 3 Ultra, costs $0.50 input and $2.20 output per million tokens.
- Canonical page: https://noometry.com/llm-pricing/nvidia
- Last updated: 2026-10-10
- Title: NVIDIA API Pricing (October 2026): Every Model per 1M Tokens

NVIDIA API prices run from $0.05 per million input tokens (Nemotron 3 Nano 30B A3B) to $0.50 (Nemotron 3 Ultra). Its highest-ranked model, Nemotron 3 Ultra, costs $0.50 input and $2.20 output per million tokens.

Last verified October 10, 2026

API prices per million tokens
|  |  |  | Cached input |  |  |  |
| --- | --- | --- | --- | --- | --- | --- |
| [Nemotron 3 Nano 30B A3B](https://noometry.com/models/nemotron-3-nano-30b-a3b) (open weights) | $0.05 | $0.20 | $0.025 | **$0.0875** | 262K | 40.6 |
| [Nemotron 3.5 Lightning](https://noometry.com/models/nemotron-3-5-lightning) (open weights) | $0.05 | $0.20 | $0.01 | **$0.0875** | 262K | 40.0 |
| [Nemotron 3 Super](https://noometry.com/models/nemotron-3-super) (open weights) | $0.08 | $0.45 | — | **$0.17** | 262K | 40.1 |
| [Nemotron 3.5 Content Safety](https://noometry.com/models/nemotron-3-5-content-safety) (open weights) | $0.20 | $0.20 | — | **$0.20** | 131K | — |
| [Nvidia Llama 3.3 Nemotron Super 49b v1.5](https://noometry.com/models/nvidia-llama-3-3-nemotron-super-49b-v1-5) (open weights) | $0.40 | $0.40 | — | **$0.40** | 131K | 40.3 |
| [Nemotron 3 Ultra](https://noometry.com/models/nemotron-3-ultra) (open weights) | $0.50 | $2.20 | $0.10 | **$0.93** | 262K | 42.5 |

## Prices on other platforms

The same models are often sold through clouds and resellers at different rates.

| Model | Route | Input | Output | Checked |
| --- | --- | --- | --- | --- |
| [Nemotron 3 Ultra](https://noometry.com/models/nemotron-3-ultra) | fireworks | $0.60 | $2.40 | 2026-10-10 |
|  | openrouter | $0.50 | $2.20 | 2026-10-10 |
|  | together | $0.60 | $3.60 | 2026-10-10 |
| [Nemotron 3 Nano 30B A3B](https://noometry.com/models/nemotron-3-nano-30b-a3b) | bedrock | $0.06 | $0.24 | 2026-10-10 |
|  | deepinfra | $0.05 | $0.20 | 2026-10-10 |
|  | openrouter | $0.06 | $0.24 | 2026-10-10 |
| [Nvidia Llama 3.3 Nemotron Super 49b v1.5](https://noometry.com/models/nvidia-llama-3-3-nemotron-super-49b-v1-5) | deepinfra | $0.40 | $0.40 | 2026-10-10 |
| [Nemotron 3 Super](https://noometry.com/models/nemotron-3-super) | bedrock | $0.15 | $0.65 | 2026-10-10 |
|  | openrouter | $0.08 | $0.45 | 2026-10-10 |
| [Nemotron 3.5 Lightning](https://noometry.com/models/nemotron-3-5-lightning) | fireworks | $0.05 | $0.20 | 2026-10-10 |
|  | openrouter | $0.07 | $0.20 | 2026-10-10 |
| [Nemotron 3.5 Content Safety](https://noometry.com/models/nemotron-3-5-content-safety) | openrouter | $0.20 | $0.20 | 2026-10-10 |

## Frequently asked questions

### How much does the NVIDIA API cost?

NVIDIA API prices run from $0.05 per million input tokens (Nemotron 3 Nano 30B A3B) to $0.50 (Nemotron 3 Ultra). Its highest-ranked model, Nemotron 3 Ultra, costs $0.50 input and $2.20 output per million tokens.

### Does NVIDIA discount cached input?

Yes. Cached input is billed at a lower rate on 3 of the 6 priced models; the cached rate is listed next to each model.
