model · class value · updated 2026-08-21
What it's for
Best for: Enterprise assistants, synthetic data, reasoning, and internal RAG.
Not ideal for: Latency-critical chat or multimodal analysis.
Role: Open / enterprise-tuned generalist
Specifications
| Provider | NVIDIA |
| API format | openai |
| Context window | 1000000 |
| Restricted access | No |
| Capabilities | Tools, Thinking |
| Execution Paths | api, self_host |
Price components
| Type | USD | Basis | Conditions | Source | Verified | Freshness |
|---|---|---|---|---|---|---|
| per_million_in | $0.09 | required · verified | OpenRouter default price-weighted routing observation; this pair matched the DeepInfra BF16 Nemotron 3 Super 120B endpoint on 2026-07-29, but the actual routed provider and billed rate can vary with availability, quantization, and routing preferences. | OpenRouter Nemotron 3 Super routing price row | 2026-08-21 | vendor · 0d |
| per_million_out | $0.40 | required · verified | OpenRouter default price-weighted routing observation; this pair matched the DeepInfra BF16 Nemotron 3 Super 120B endpoint on 2026-07-29, but the actual routed provider and billed rate can vary with availability, quantization, and routing preferences. | OpenRouter Nemotron 3 Super routing price row | 2026-08-21 | vendor · 0d |
Cost at three usage tiers
| A few questions and small tasks each day. | $0.87/mo |
| Daily use, plus some scheduled tasks. | $2.90/mo |
| Working for you most of the day, including web tasks. | $8.70/mo |
Token cost only; add a host for the full-stack estimate.
Compare