Sarvam: Sarvam 105B (Free)
Sarvam
Feb 18, 2026
sarvam/sarvam-105b:free
Context Length128,000
Input Price-
Output Price-
Max Output
Blended Price-

Sarvam 105B is Sarvam AI's flagship open-source Mixture-of-Experts (MoE) large language model, trained from scratch as India's first competitive 105-billion-parameter model. It delivers strong performance across Indian languages and excels in enterprise-grade applications, with capabilities in multi-step reasoning, mathematics, coding, knowledge retrieval, and instruction-following.

Input Modalitiestext
Output Modalitiestext
Modalitytext->text
Provider Details

sarvam

Input Cost-
Output Cost-
Context Length128,000
Max Output128,000
Latency-
Throughput-
Supported Parameters
Parameter
Core (Sampling & basic generation)
max_tokensMaximum tokens to generate.
nNumber of completions to generate.
seedSeed for deterministic sampling.
stopSequences where generation stops.
temperatureControls randomness of output.
top_pNucleus sampling threshold.
Penalties (Token Penalties)
frequency_penaltyPenalizes frequent tokens.
presence_penaltyPenalizes tokens already present.
Reasoning (Reasoning Controls)
include_reasoningInclude reasoning trace in the response.
reasoningProvider-specific reasoning configuration.

Throughput

Latency

Make Your First API Call

Get started by making your first request to FastRouter. Use your favorite SDK or make direct API calls.

Copy
from openai import OpenAI
client = OpenAI(
base_url="https://api.fastrouter.ai/api/v1", # FastRouter base URL
api_key= "<FASTROUTER_API_KEY>", # Replace with your FastRouter API key
)
completion = client.chat.completions.create(
model="sarvam/sarvam-105b:free", # Replace with your model ID
messages=[
{ "role": "user", "content": "What is the meaning of life?" }
]
)
print(completion.choices[0].message.content)

Before you run this code:

  1. Replace <FASTROUTER_API_KEY> with your actual API key
  2. Install the OpenAI SDK: pip install openai
  3. Run the code, then check your dashboard for analytics
Frequently Asked Questions

More models from Sarvam
sarvam/bulbul:v2

Sarvam: Bulbul V2 is a text-to-speech (TTS) model from Sarvam AI, delivering real-time, natural-sounding speech in 11 Indian languages at a fraction of global costs. It builds on prior versions with key enhancements including improved audio quality, 30+ speaker voices (expanded from six distinct voices in v1 for diverse use cases like professional or conversational tones), and adjustable speech speed from 0.5x to 2.0x for customized delivery.

sarvam/saaras:v3

Saaras V3 is Sarvam AI's next-generation automatic speech recognition (ASR) model engineered for Indian languages and real-world speech conditions. Saaras V3 supports 23 languages (22 official Indian languages plus English) within a unified multilingual model. The model achieves approximately 19% Word Error Rate (WER) on the IndicVoices benchmark, improving significantly from its predecessor Saaras V2.5 which had ~22% WER. On the 10 most popular languages in the IndicVoices dataset, it achieves 19.3% WER, and performance advantages widen for lower-resource Indian languages

sarvam/sarvam-105b

Sarvam 105B is Sarvam AI's flagship open-source Mixture-of-Experts (MoE) large language model (LLM), trained from scratch as India's first competitive 105-billion-parameter model. It delivers state-of-the-art performance across Indian languages and excels in enterprise-grade applications, with strong capabilities in multi-step reasoning, mathematics, coding, knowledge retrieval, and instruction-following.

Model scores sourced from ArtificialAnalysis