NVIDIA logo

NVIDIA AI models

2 NVIDIA models on FastRouter, all behind one OpenAI-compatible API. Compare pricing, context windows and benchmarks, then open any model for its providers and code samples.

Filter in catalog
Models
2
Largest context
1.05M
Nemotron 3 Nano 30B (Free)

All NVIDIA models

NVIDIA logo

NVIDIA Nemotron 3 Nano 30B is a hybrid Mamba-2/MoE/Attention model with 3.5B active parameters, trained from scratch as a unified model for both reasoning and non-reasoning tasks. Reasoning traces can be toggled via a chat template flag, and the model supports tool use within a 1M-token context window, making it well suited for specialized agentic applications.

nvidia/nemotron-3-nano-30b:freeJul 9, 2026
Context
1.05M
Price /1M
Free
Intel
—
NVIDIA logo

NVIDIA Nemotron 3 Super is a 120B-parameter open mixture-of-experts model that activates just 12B parameters per token, delivering high compute efficiency and strong reasoning, agentic, and conversational capabilities for complex multi-agent applications. It supports tool use, thinking/reasoning traces, and a large 262K-token context window, and is optimized for collaborative agents and high-volume workloads such as IT ticket automation.

nvidia/nemotron-3-super:freeJul 9, 2026
Context
262K
Price /1M
Free
Intel
—

Frequently asked questions

Model scores sourced from ArtificialAnalysis

NVIDIA AI Models: API Pricing & Benchmarks | FastRouter.ai