DeepSeek logo

DeepSeek AI models

5 DeepSeek models on FastRouter, all behind one OpenAI-compatible API. Compare pricing, context windows and benchmarks, then open any model for its providers and code samples.

Filter in catalog
Models
5
Largest context
1.05M
DeepSeek V4.1 Flash
Lowest input /1M
$0.09
V4 Flash
Top intelligence
39.5
DeepSeek V4.1 Flash

All DeepSeek models

DeepSeek logo

DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with a 552B-parameter backbone, activating just 8B parameters per token during prefill and 16B during decoding. It natively processes images and text, supports a 1M-token context window, and introduces a Causal Encoder-Decoder architecture with Compressed Sparse Attention 2 to make long-context inference more efficient. Configurable reasoning-effort levels let developers trade latency and cost for deeper deliberation. Compared with DeepSeek-V4-Flash-0731, it reduces the global KV-cache footprint by roughly 4x while delivering stronger coding, reasoning, and agentic performance.

deepseek/deepseek-v4.1-flashSep 14, 2026
Context
1.05M
Price /1M
$0.20 in$0.60 out
Intel
39.5
DeepSeek logo

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts large language model from DeepSeek with a 1M-token context window and strong general, coding, and reasoning performance. It is designed as a fast, low-cost frontier model suitable for long-context chat, agents, and tools. The Flash variant emphasizes high throughput and low latency while maintaining competitive reasoning quality.

deepseek/deepseek-v4-flashJul 31, 20267.9s latency
Context
1.05M
Price /1M
$0.09 in$0.18 out
Intel
34.3
DeepSeek logo

DeepSeek-V4-Pro is a 1.6 trillion parameter Mixture-of-Experts (MoE) language model from DeepSeek AI, with 49 billion parameters activated per token and support for a 1 million token context window. It features a hybrid attention architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA), which reduces inference FLOPs to 27% and KV cache to 10% of DeepSeek-V3.2 at 1M context, enabling 83.5% MRCR comprehension. The model was pre-trained on over 32 trillion tokens using the Muon optimizer and a two-stage post-training process involving SFT, RL with GRPO, and on-policy distillation for domain-specific expertise.

deepseek/deepseek-v4-proApr 24, 20264.9s latency
Context
1.05M
Price /1M
$1.74 in$3.48 out
Intel
36.0
DeepSeek logo

DeepSeek-V3.2 is positioned as a next-generation “general-purpose + reasoning” model intended to be a daily driver at roughly frontier (GPT‑5-class) performance on broad tasks like coding, math, agents, and general chat. It is released in several variants (such as V3.2, V3.2-Exp, and V3.2-Speciale), sharing the same core architecture but targeting different trade‑offs between efficiency and maximum reasoning power.

deepseek/deepseek-v3.2Dec 1, 202541s latency
Context
164K
Price /1M
$0.26 in$0.38 out
Intel
21.5
DeepSeek logo

DeepSeek V3.1 is an advanced open-source hybrid AI model designed to balance powerful reasoning with high-speed efficiency. It uniquely supports two inference modes—"Think" for deep reasoning and "Non-Think" for direct, lightweight tasks—making it versatile across use cases. Built with a mixture-of-experts architecture, it scales to 685B parameters while activating only 37B per token, enabling cost-effective performance. With a 128K context window, it can handle large documents, codebases, and complex multi-step workflows. DeepSeek V3.1 is optimized for tool use, agent-based applications, and enterprise deployment through open weights and developer-friendly APIs.

deepseek/deepseek-v3.1Aug 21, 202516s latency
Context
164K
Price /1M
$0.25 in$0.95 out
Intel
13.5

Frequently asked questions

Model scores sourced from ArtificialAnalysis

DeepSeek AI Models: API Pricing & Benchmarks | FastRouter.ai