NVIDIA: Nemotron 3 Super 120B (Free)
NVIDIA
Jul 9, 2026
nvidia/nemotron-3-super:free
Context Length262,144
Input Price-
Output Price-
Max Output
Blended Price-

NVIDIA Nemotron 3 Super is a 120B-parameter open mixture-of-experts model that activates just 12B parameters per token, delivering high compute efficiency and strong reasoning, agentic, and conversational capabilities for complex multi-agent applications. It supports tool use, thinking/reasoning traces, and a large 262K-token context window, and is optimized for collaborative agents and high-volume workloads such as IT ticket automation.

Input Modalitiestext
Output Modalitiestext
Modalitytext->text
TokenizerOther
Provider Details

ollama

JSON
Input Cost-
Output Cost-
Context Length262,144
Max Output32,768
Latency-
Throughput-
Supported Parameters
Parameter
Core (Sampling & basic generation)
max_completion_tokensMaximum tokens in the completion.
seedSeed for deterministic sampling.
stopSequences where generation stops.
temperatureControls randomness of output.
top_kLimits sampling to top K tokens.
top_pNucleus sampling threshold.
Penalties (Token Penalties)
frequency_penaltyPenalizes frequent tokens.
presence_penaltyPenalizes tokens already present.
Reasoning (Reasoning Controls)
reasoningProvider-specific reasoning configuration.
Tool Use (Function Calls & Tools)
tool_choiceHow the model should use tools.

Throughput

Latency

Make Your First API Call

Get started by making your first request to FastRouter. Use your favorite SDK or make direct API calls.

Copy
from openai import OpenAI
client = OpenAI(
base_url="https://api.fastrouter.ai/api/v1", # FastRouter base URL
api_key= "<FASTROUTER_API_KEY>", # Replace with your FastRouter API key
)
completion = client.chat.completions.create(
model="nvidia/nemotron-3-super:free", # Replace with your model ID
messages=[
{ "role": "user", "content": "What is the meaning of life?" }
]
)
print(completion.choices[0].message.content)

Before you run this code:

  1. Replace <FASTROUTER_API_KEY> with your actual API key
  2. Install the OpenAI SDK: pip install openai
  3. Run the code, then check your dashboard for analytics
Frequently Asked Questions

More models from NVIDIA
nvidia/nemotron-3-nano-30b:free

NVIDIA Nemotron 3 Nano 30B is a hybrid Mamba-2/MoE/Attention model with 3.5B active parameters, trained from scratch as a unified model for both reasoning and non-reasoning tasks. Reasoning traces can be toggled via a chat template flag, and the model supports tool use within a 1M-token context window, making it well suited for specialized agentic applications.

Model scores sourced from ArtificialAnalysis