Google: Gemma 4 26B (Free)
Google
Jul 9, 2026
google/gemma4-26b:free
Context Length262,144
Input Price-
Output Price-
Max Output
Blended Price-

Gemma 4 26B is a 31B-parameter dense multimodal model from Google DeepMind, handling text, image, and audio input to generate text output. It supports an explicit thinking mode via control tokens, tool use, and a 256K-token context window, and is released under an Apache 2.0 license.

Input Modalitiestext, image
Output Modalitiestext
Modalitytext+image->text
TokenizerOther
Provider Details

ollama

JSON
Input Cost-
Output Cost-
Context Length262,144
Max Output32,768
Latency-
Throughput-
Supported Parameters
Parameter
Core (Sampling & basic generation)
max_completion_tokensMaximum tokens in the completion.
seedSeed for deterministic sampling.
stopSequences where generation stops.
temperatureControls randomness of output.
top_kLimits sampling to top K tokens.
top_pNucleus sampling threshold.
Penalties (Token Penalties)
frequency_penaltyPenalizes frequent tokens.
presence_penaltyPenalizes tokens already present.
Reasoning (Reasoning Controls)
reasoningProvider-specific reasoning configuration.
Tool Use (Function Calls & Tools)
tool_choiceHow the model should use tools.

Throughput

Latency

Make Your First API Call

Get started by making your first request to FastRouter. Use your favorite SDK or make direct API calls.

Copy
from openai import OpenAI
client = OpenAI(
base_url="https://api.fastrouter.ai/api/v1", # FastRouter base URL
api_key= "<FASTROUTER_API_KEY>", # Replace with your FastRouter API key
)
completion = client.chat.completions.create(
model="google/gemma4-26b:free", # Replace with your model ID
messages=[
{ "role": "user", "content": "What is the meaning of life?" }
]
)
print(completion.choices[0].message.content)

Before you run this code:

  1. Replace <FASTROUTER_API_KEY> with your actual API key
  2. Install the OpenAI SDK: pip install openai
  3. Run the code, then check your dashboard for analytics
Frequently Asked Questions

More models from Google
google/gemini-3.8-flash

Gemini 3.8 Flash is Google’s most capable Flash-series multimodal model, optimized for long-horizon software engineering, autonomous agents, and complex enterprise workflows while maintaining high speed and cost efficiency. It supports very long contexts (up to 1,048,576 tokens) across text, images, video, audio, and PDFs, and offers advanced capabilities like tool calling, computer use, file search, and structured outputs. It is suitable for agentic applications, complex reasoning, and high-throughput production workloads.

google/gemini-3.1-flash-image-preview

Google's gemini-3.1-flash-image-preview (codenamed Nano Banana 2) is a preview version of the Gemini 3.1 Flash Image model, optimized for high-speed image generation and editing with Pro-level quality, balancing performance, low latency, and cost efficiency.

google/gemini-3.7-flash

Gemini 3.7 Flash is a highly capable, natively multimodal Gemini 3-series model optimized for fast, agentic workflows and complex reasoning across text, images, video, audio, and PDFs. It supports large 1M-token contexts, advanced tools like function calling, file search, computer use, and structured outputs, and configurable "thinking" levels for more deliberate reasoning. Designed for production applications via the Gemini API, it balances speed, cost-efficiency, and strong multimodal understanding.

google/gemini-3.6-flash

Gemini 3.6 Flash is a multimodal frontier-intelligence model optimized for speed and cost, supporting text, image, video, audio, and PDF inputs with text-only outputs. It targets complex agentic workflows, strong code generation, and spatial/multimodal reasoning while maintaining a 1M-token context window and large output capacity. The model supports advanced tools such as function calling, computer use, file search, search grounding, structured outputs, and thinking.

google/gemini-3.5-flash-lite

Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution in agentic workflows and document parsing. It supports text, image, video, audio, and PDF inputs with text-only outputs, a 1M-token context window, and up to 65,536 output tokens. The model is designed for subagent tasks, simple data extraction, and applications where latency and API cost are primary constraints.

Google Gemini 2.5 Flash Image (often nicknamed "nano-banana") is Google's latest state-of-the-art multimodal model for image generation and editing, designed for both developers and enterprises. It enables users to create, blend, and edit images with natural language prompts and multi-image input, all while benefiting from Gemini's deep real-world understanding and conversational editing features.

google/gemini-2.5-flash

Gemini 2.5 Flash is Google's hybrid reasoning model offering an optimal balance between performance, cost, and latency. It features full thinking capabilities that can be toggled on/off or fine-tuned using a thinking budget parameter (0-24576 tokens). This model excels at complex tasks requiring multi-step reasoning while maintaining the efficiency needed for high-volume and real-time applications. It outperforms comparable models in reasoning benchmarks while maintaining lower cost and latency profiles, making it ideal for enterprise deployments requiring both deep analysis and scalability.

google/gemini-2.5-pro

Gemini 2.5 Pro is Google's premiere thinking model built for enterprise-grade applications requiring sophisticated reasoning. It delivers top performance on critical benchmarks, featuring built-in reasoning capabilities that enable it to plan and analyze before responding to complex queries. The model maintains high accuracy across STEM domains with native multimodal processing and an extensive context window. This preview release allows early access to Google's most advanced model, which excels in code development, scientific problem-solving, and multi-step logical tasks while maintaining human-aligned responses. The preview designation indicates ongoing optimization before general availability.

google/gemini-3-flash-preview

Google's gemini-3-flash-preview is a high-speed, cost-effective AI model from the Gemini 3 series, designed for agentic workflows, multi-turn chats, and coding tasks. It delivers near-Pro-level reasoning and tool use with lower latency than larger variants, supporting a 1M token context window and multimodal inputs like text, images, audio, video, and PDFs.

Gemini 3 Pro Image Preview is a high-end image generation and editing model in the Gemini 3 family (also known as Nano Banana Pro), built to produce studio-quality visuals with strong reasoning over complex prompts. It is optimized for detailed, multi-step creative workflows where you need both visual fidelity and precise control over content, layout, and text inside images.

Gemini 3.1 Flash-Lite Image is a cost-efficient multimodal Gemini 3.1 variant optimized for high-volume image workflows, supporting text and image inputs with both image and text outputs. It is suitable for image generation, editing, and related multimodal tasks where low latency and price are important. The model is available via the Gemini API in Google AI Studio and for enterprises via Vertex AI.

google/gemini-3.1-pro-preview

Google's gemini-3.1-pro-preview (also called Gemini 3.1 Pro Preview) is a preview-stage, multimodal AI model from Google DeepMind, optimized for advanced reasoning, agentic workflows, software engineering, and complex problem-solving across text, images, video, audio, and PDFs. It refines the Gemini 3 Pro series with improved thinking, token efficiency, factual consistency, and reliability for multi-step tasks, tool use, and long-horizon stability. Key specs include a 1M input token context window (1,048,576 tokens), 64K-65K output tokens, and support for capabilities like function calling, structured outputs, code execution, search grounding, and a new "MEDIUM" thinking level for balancing cost, speed, and performance.

google/gemini-3.5-flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model optimized for speed, coding, and agentic workflows. It supports text, image, video, audio, and PDF inputs with text output, and provides strong reasoning/tool-use at lower latency.

google/gemini-embedding-001

State-of-the-art Gemini embedding model (3072-dim) optimised for retrieval and semantic similarity.

google/gemini-omni-1.1-flash

Gemini Omni 1.1 Flash is Google's high-performance multimodal video generation and editing model, built on native multimodality (processing text, image, audio, and video simultaneously), conversational editing via the Interactions API, and world knowledge combining physics understanding with Gemini's broader knowledge. It generates 3-10 second video clips with audio at up to 4K (upscaled), supports text-to-video, image-to-video, first/last frame interpolation, subject and video reference conditioning, and stateful multi-turn editing and extension up to a cumulative 40 seconds. All outputs carry SynthID watermarking.

google/gemini-omni-flash-preview

Gemini Omni Flash is a preview model designed for fast, conversational video generation and editing. It turns text, images, video, and audio into 720p video (3–10 seconds, 24 FPS) and supports iterative refinement through the Interactions API. Input modalities include text, image, video (up to 10s for editing), and audio. Tasks: text_to_video, image_to_video, reference_to_video, and edit. Aspect ratios: 16:9 (default) and 9:16. Delivery: base64 or uri. No free tier.

google/gemma-4-26b-a4b-it

Gemma 4 26B A4B IT is a 26B-parameter Mixture-of-Experts, instruction-tuned, open-weight multimodal model from Google DeepMind that processes text, image, and video inputs to produce text outputs. It targets long-context (up to ~256K tokens) reasoning, coding, and chat use cases while remaining efficient via 4B active parameters. The model is suitable for agents, RAG, and tool-calling style applications where high quality and large context are required.

google/gemma-4-31b-it

google/gemma-4-31b-it is Google DeepMind’s instruction-tuned 31B Gemma 4 model, built for strong reasoning, coding, agentic workflows, and multimodal understanding. It supports text and image input, has a 256K context window, and is positioned as the dense, higher-quality counterpart to the 26B MoE variant.

google/veo2

google/veo2 is Google DeepMind’s AI video generation model for creating high-quality, realistic videos from text prompts or image references. It’s built for strong motion quality, cinematic detail, and accurate prompt following, making it useful for storyboarding, concept visualization, and short-form video generation.

google/veo3

google/veo3 is Google DeepMind’s advanced AI video generation model that creates high-quality, cinematic videos with synchronized audio from text prompts or image references. It’s designed for realistic motion, strong prompt adherence, and audiovisual consistency, making it useful for short-form storytelling, concept visualization, and professional content creation.

google/veo3-fast

google/veo3-fast is Google’s speed-optimized AI video generation model for creating high-quality, cinematic videos with native audio from text or image prompts. It’s designed for rapid iteration, fast prototyping, and production workflows where turnaround time matters, while still preserving strong visual fidelity and prompt adherence.

Google Veo 3.1 is a preview model code in the Gemini API for Google's Veo 3.1, a state-of-the-art cinematic video generation engine optimized for professional-grade 4K output, natively synchronized audio, and complex camera movements with high temporal consistency.

google/veo3.1-fast

Google Veo 3.1 Fast is a generally available (GA) version of Google's Veo 3.1 Fast video generation model in Vertex AI and Gemini API, optimized for faster inference while delivering high-quality videos with synchronized native audio from text or image prompts.

google/veo3.1-lite

Google Veo 3.1 Lite, a high-efficiency, cost-effective video generation model in the Gemini API and Vertex AI, designed for developers building high-volume applications at under 50% the cost of Veo 3.1 Fast with equivalent speed.

Model scores sourced from ArtificialAnalysis