DeepSeek V4.1 Flash
Specifications
- Input
- Output
- Context window
- 1M tokens
- Veröffentlicht
- Sep 2026
Performance
- Speed
- 183 t/s
- TTFT
- —
- Latency
- 194 ms
- Intelligence
- —
Pricing
- Eingabe
- €0.20 per 1M tokens
- Ausgabe
- €1.00 per 1M tokens
€0.01 Cache-Hit
Über dieses Modell
DeepSeek V4.1 Flash is a 552B-parameter multimodal Mixture-of-Experts (MoE) chat model from DeepSeek AI, activating only 8B parameters during prefill and 16B during decode thanks to its novel Causal Encoder-Decoder (CED) architecture with Compressed Sparse Attention 2 (CSA2). The model natively processes images and text, supports a 1M-token context window, and features continuously controllable reasoning effort (1–100) for trading inference cost against accuracy. It achieves 90.9 on GPQA Diamond, 90.6 on Terminal-Bench 2.1, and 74.2 on DeepSWE v1.1, demonstrating frontier-level agentic and coding performance. The model is released under the MIT License.
Technische Daten
- Fähigkeiten
- Eingabe-Modalitäten
- Ausgabe-Modalitäten
- Reasoning
- Hybrid Standard on
Knowledge horizon
Veröffentlicht Sep 2026
Today
Since release 0 mo
See also
Modell zum Vergleich hinzufügen
Nach einem Modell suchen