DeepSeek V4 Flash 0731
Specifications
- Input
- Output
- Context window
- 1M tokens
- Released
- Jul 2026
Performance
- Speed
- 125 t/s
- TTFT
- 329 ms
- Latency
- 281 ms
- Intelligence
- —
Pricing
- Input
- €0.10 per 1M tokens
- Output
- €0.30 per 1M tokens
About this model
DeepSeek V4 Flash 0731 is the official release of DeepSeek V4 Flash, superseding the preview version. It is a 304B-parameter Mixture-of-Experts model (256 routed experts plus one shared expert, 6 active per token) with an attached DSpark speculative-decoding module and a 1M-token context window. Deliberation is controlled through three reasoning effort levels — low, high and max — with up to 384K output tokens recommended at the high and max levels. On published agentic benchmarks it reaches 82.7 on Terminal Bench 2.1, 70.3 on Toolathlon-Verified and 54.4 on DeepSWE, outperforming DeepSeek V4 Pro (Preview) despite a far smaller activated parameter count. Text-only, with function calling and structured output, released under the MIT License.
Technical specifications
- Capabilities
- Input modalities
- Output modalities
- Reasoning
- Hybrid Default on