Google logo

Nano Banana 2.1 (Gemini Nano Banana 2.1)

NEW
GoogleReleased Oct 6, 2026google/gemini-nano-banana-2.1
Compare

Google's gemini-nano-banana-2.1 (Nano Banana 2.1) is the successor to Nano Banana 2 (Gemini 3.1 Flash Image), built on Gemini 3.6 Flash for high-efficiency image generation and conversational editing at Flash-level speed and cost. It improves visual quality, text rendering, infographic layout accuracy, and multi-turn character consistency across 1K, 2K, and 4K output, fixes tiling artifacts on wide and panoramic aspect ratios, supports up to 14 reference images (up to 4 characters and 10 objects), grounding with Google Web and Image Search, and configurable thinking levels (minimal, medium, high). All outputs carry a SynthID watermark.

Context
131K
Max output
33K
Input /1M
$1.50
Output /1M
$0.0336/img
AcceptsImageTextProducesImageTextTokenizer Gemini

Providers

FastRouter routes your requests to this provider. You can also bring your own key.

Pricing tier
Google AI Studio logo
Google AI Studio
googleaistudio
Input /1M$1.50
Output /1M$0.0336/img
Context131K
Max output33K
Supports tool calling

Supported parameters

Request parameters you can send with this model.

Parameter
Type
Google AI Studio
Core (Sampling & basic generation)
candidateCountNumber of completions to generate.
number
1 to 8default: 1
maxOutputTokensMaximum tokens to generate.
number
1 to 32768
seedSeed for deterministic sampling.
number
-
stopSequencesSequences where generation stops.
array
max 5 items
temperatureControls randomness of output.
number
0 to 2
topKLimits sampling to top K tokens.
number
≥ 1
topPNucleus sampling threshold.
number
0 to 1
Reasoning (Reasoning Controls)
thinkingConfigProvider-specific reasoning configuration.
object
Tool Use (Function Calls & Tools)
toolsTool/function definitions available to the model.
array
Output & Format (Response Formatting)
responseFormatStructured output format configuration.
object

Make your first API call

OpenAI-compatible. Point your SDK at FastRouter, or call the API directly.

Full API docs
Replace <FASTROUTER_API_KEY> with your key · Get your key →Install
OpenAI SDKHTTP
import base64
from openai import OpenAI
client = OpenAI(
api_key="<FASTROUTER_API_KEY>",
base_url="https://api.fastrouter.ai/api/v1"
)
img = client.images.generate(
model="google/gemini-nano-banana-2.1",
prompt="A cute baby sea otter",
n=1,
size="1024x1024"
)
image_base64 = img.data[0].b64_json
image_bytes = base64.b64decode(image_base64)
with open("otter.png", "wb") as f:
f.write(image_bytes)

Frequently asked questions

More models from Google

Google logo
Nano Banana 2 (Gemini 3.1 Flash Image Preview)google/gemini-3.1-flash-image-preview

Google's gemini-3.1-flash-image-preview (codenamed Nano Banana 2) is a preview version of the Gemini 3.1 Flash Image model, optimized for high-speed image generation and editing with Pro-level quality, balancing performance, low latency, and cost efficiency.

Google logo
Gemini 3.8 Flashgoogle/gemini-3.8-flash

Gemini 3.8 Flash is Google’s most capable Flash-series multimodal model, optimized for long-horizon software engineering, autonomous agents, and complex enterprise workflows while maintaining high speed and cost efficiency. It supports very long contexts (up to 1,048,576 tokens) across text, images, video, audio, and PDFs, and offers advanced capabilities like tool calling, computer use, file search, and structured outputs. It is suitable for agentic applications, complex reasoning, and high-throughput production workloads.

Google logo
Gemini 3.7 Flashgoogle/gemini-3.7-flash

Gemini 3.7 Flash is a highly capable, natively multimodal Gemini 3-series model optimized for fast, agentic workflows and complex reasoning across text, images, video, audio, and PDFs. It supports large 1M-token contexts, advanced tools like function calling, file search, computer use, and structured outputs, and configurable "thinking" levels for more deliberate reasoning. Designed for production applications via the Gemini API, it balances speed, cost-efficiency, and strong multimodal understanding.

Google logo
Gemini 3.6 Flashgoogle/gemini-3.6-flash

Gemini 3.6 Flash is a multimodal frontier-intelligence model optimized for speed and cost, supporting text, image, video, audio, and PDF inputs with text-only outputs. It targets complex agentic workflows, strong code generation, and spatial/multimodal reasoning while maintaining a 1M-token context window and large output capacity. The model supports advanced tools such as function calling, computer use, file search, search grounding, structured outputs, and thinking.

Google logo
Gemini 2.5 Flash Image (Nano Banana)google/gemini-2.5-flash-image

Google Gemini 2.5 Flash Image (often nicknamed "nano-banana") is Google's latest state-of-the-art multimodal model for image generation and editing, designed for both developers and enterprises. It enables users to create, blend, and edit images with natural language prompts and multi-image input, all while benefiting from Gemini's deep real-world understanding and conversational editing features.

Google logo
Gemini 2.5 Flashgoogle/gemini-2.5-flash

Gemini 2.5 Flash is Google's hybrid reasoning model offering an optimal balance between performance, cost, and latency. It features full thinking capabilities that can be toggled on/off or fine-tuned using a thinking budget parameter (0-24576 tokens). This model excels at complex tasks requiring multi-step reasoning while maintaining the efficiency needed for high-volume and real-time applications. It outperforms comparable models in reasoning benchmarks while maintaining lower cost and latency profiles, making it ideal for enterprise deployments requiring both deep analysis and scalability.

Google logo
Gemini 2.5 Pro google/gemini-2.5-pro

Gemini 2.5 Pro is Google's premiere thinking model built for enterprise-grade applications requiring sophisticated reasoning. It delivers top performance on critical benchmarks, featuring built-in reasoning capabilities that enable it to plan and analyze before responding to complex queries. The model maintains high accuracy across STEM domains with native multimodal processing and an extensive context window. This preview release allows early access to Google's most advanced model, which excels in code development, scientific problem-solving, and multi-step logical tasks while maintaining human-aligned responses. The preview designation indicates ongoing optimization before general availability.

Google logo
Gemini 3 Flash Previewgoogle/gemini-3-flash-preview

Google's gemini-3-flash-preview is a high-speed, cost-effective AI model from the Gemini 3 series, designed for agentic workflows, multi-turn chats, and coding tasks. It delivers near-Pro-level reasoning and tool use with lower latency than larger variants, supporting a 1M token context window and multimodal inputs like text, images, audio, video, and PDFs.

Google logo
Gemini 3 Pro Image Preview (Nano Banana Pro)google/gemini-3-pro-image-preview

Gemini 3 Pro Image Preview is a high-end image generation and editing model in the Gemini 3 family (also known as Nano Banana Pro), built to produce studio-quality visuals with strong reasoning over complex prompts. It is optimized for detailed, multi-step creative workflows where you need both visual fidelity and precise control over content, layout, and text inside images.

Google logo
Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image)google/gemini-3.1-flash-lite-image

Gemini 3.1 Flash-Lite Image is a cost-efficient multimodal Gemini 3.1 variant optimized for high-volume image workflows, supporting text and image inputs with both image and text outputs. It is suitable for image generation, editing, and related multimodal tasks where low latency and price are important. The model is available via the Gemini API in Google AI Studio and for enterprises via Vertex AI.

Google logo
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview

Google's gemini-3.1-pro-preview (also called Gemini 3.1 Pro Preview) is a preview-stage, multimodal AI model from Google DeepMind, optimized for advanced reasoning, agentic workflows, software engineering, and complex problem-solving across text, images, video, audio, and PDFs. It refines the Gemini 3 Pro series with improved thinking, token efficiency, factual consistency, and reliability for multi-step tasks, tool use, and long-horizon stability. Key specs include a 1M input token context window (1,048,576 tokens), 64K-65K output tokens, and support for capabilities like function calling, structured outputs, code execution, search grounding, and a new "MEDIUM" thinking level for balancing cost, speed, and performance.

Google logo
Gemini 3.5 Flashgoogle/gemini-3.5-flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model optimized for speed, coding, and agentic workflows. It supports text, image, video, audio, and PDF inputs with text output, and provides strong reasoning/tool-use at lower latency.

Google logo
Gemini 3.5 Flash-Litegoogle/gemini-3.5-flash-lite

Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution in agentic workflows and document parsing. It supports text, image, video, audio, and PDF inputs with text-only outputs, a 1M-token context window, and up to 65,536 output tokens. The model is designed for subagent tasks, simple data extraction, and applications where latency and API cost are primary constraints.

Google logo
gemini-embeddinggoogle/gemini-embedding-001

State-of-the-art Gemini embedding model (3072-dim) optimised for retrieval and semantic similarity.

Google logo
Gemini Omni 1.1 Flashgoogle/gemini-omni-1.1-flash

Gemini Omni 1.1 Flash is Google's high-performance multimodal video generation and editing model, built on native multimodality (processing text, image, audio, and video simultaneously), conversational editing via the Interactions API, and world knowledge combining physics understanding with Gemini's broader knowledge. It generates 3-10 second video clips with audio at up to 4K (upscaled), supports text-to-video, image-to-video, first/last frame interpolation, subject and video reference conditioning, and stateful multi-turn editing and extension up to a cumulative 40 seconds. All outputs carry SynthID watermarking.

Google logo
Gemini Omni Flash Previewgoogle/gemini-omni-flash-preview

Gemini Omni Flash is a preview model designed for fast, conversational video generation and editing. It turns text, images, video, and audio into 720p video (3–10 seconds, 24 FPS) and supports iterative refinement through the Interactions API. Input modalities include text, image, video (up to 10s for editing), and audio. Tasks: text_to_video, image_to_video, reference_to_video, and edit. Aspect ratios: 16:9 (default) and 9:16. Delivery: base64 or uri. No free tier.

Google logo
Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it

Gemma 4 26B A4B IT is a 26B-parameter Mixture-of-Experts, instruction-tuned, open-weight multimodal model from Google DeepMind that processes text, image, and video inputs to produce text outputs. It targets long-context (up to ~256K tokens) reasoning, coding, and chat use cases while remaining efficient via 4B active parameters. The model is suitable for agents, RAG, and tool-calling style applications where high quality and large context are required.

Google logo
Gemma 4 31Bgoogle/gemma-4-31b-it

google/gemma-4-31b-it is Google DeepMind’s instruction-tuned 31B Gemma 4 model, built for strong reasoning, coding, agentic workflows, and multimodal understanding. It supports text and image input, has a 256K context window, and is positioned as the dense, higher-quality counterpart to the 26B MoE variant.

Google logo
Gemma 4 26B (Free)google/gemma4-26b:free

Gemma 4 26B is a 31B-parameter dense multimodal model from Google DeepMind, handling text, image, and audio input to generate text output. It supports an explicit thinking mode via control tokens, tool use, and a 256K-token context window, and is released under an Apache 2.0 license.

Google logo
Veo 2google/veo2

google/veo2 is Google DeepMind’s AI video generation model for creating high-quality, realistic videos from text prompts or image references. It’s built for strong motion quality, cinematic detail, and accurate prompt following, making it useful for storyboarding, concept visualization, and short-form video generation.

Google logo
Veo 3google/veo3

google/veo3 is Google DeepMind’s advanced AI video generation model that creates high-quality, cinematic videos with synchronized audio from text prompts or image references. It’s designed for realistic motion, strong prompt adherence, and audiovisual consistency, making it useful for short-form storytelling, concept visualization, and professional content creation.

Google logo
Veo 3 Fastgoogle/veo3-fast

google/veo3-fast is Google’s speed-optimized AI video generation model for creating high-quality, cinematic videos with native audio from text or image prompts. It’s designed for rapid iteration, fast prototyping, and production workflows where turnaround time matters, while still preserving strong visual fidelity and prompt adherence.

Google logo
Google Veo 3.1google/veo3.1

Google Veo 3.1 is a preview model code in the Gemini API for Google's Veo 3.1, a state-of-the-art cinematic video generation engine optimized for professional-grade 4K output, natively synchronized audio, and complex camera movements with high temporal consistency.

Google logo
Google Veo 3.1 Fastgoogle/veo3.1-fast

Google Veo 3.1 Fast is a generally available (GA) version of Google's Veo 3.1 Fast video generation model in Vertex AI and Gemini API, optimized for faster inference while delivering high-quality videos with synchronized native audio from text or image prompts.

Google logo
Google Veo 3.1 Litegoogle/veo3.1-lite

Google Veo 3.1 Lite, a high-efficiency, cost-effective video generation model in the Gemini API and Vertex AI, designed for developers building high-volume applications at under 50% the cost of Veo 3.1 Fast with equivalent speed.

Model scores sourced from ArtificialAnalysis