DeepSeek V4.1 Flash
Specifications
- Input
- Output
- Context window
- 1M tokens
- Released
- Sep 2026
Performance
- Speed
- 184 t/s
- TTFT
- 564 ms
- Latency
- 147 ms
- Intelligence
- —
Pricing
- Input
- €0.20 per 1M tokens
- Output
- €1.00 per 1M tokens
€0.01 cache hit
About this model
DeepSeek V4.1 Flash is a 552B-parameter multimodal Mixture-of-Experts (MoE) chat model from DeepSeek AI, activating only 8B parameters during prefill and 16B during decode thanks to its novel Causal Encoder-Decoder (CED) architecture with Compressed Sparse Attention 2 (CSA2). The model natively processes images and text, supports a 1M-token context window, and features continuously controllable reasoning effort (1–100) for trading inference cost against accuracy. It achieves 90.9 on GPQA Diamond, 90.6 on Terminal-Bench 2.1, and 74.2 on DeepSWE v1.1, demonstrating frontier-level agentic and coding performance. The model is released under the MIT License.
Technical specifications
- Capabilities
- Input modalities
- Output modalities
- Reasoning
- Hybrid Default on
Knowledge horizon
Released Sep 2026
Today
Since release 0 mo
See also
Add Model to Comparison
Search for a model to add