DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with a 552B-parameter backbone, activating just 8B parameters per token during prefill and 16B during decoding. It natively processes images and text, supports a 1M-token context window, and introduces a Causal Encoder-Decoder architecture with Compressed Sparse Attention 2 to make long-context inference more efficient. Configurable reasoning-effort levels let developers trade latency and cost for deeper deliberation. Compared with DeepSeek-V4-Flash-0731, it reduces the global KV-cache footprint by roughly 4x while delivering stronger coding, reasoning, and agentic performance.
deepseek/deepseek-v4.1-flashSep 14, 2026