Qwen3.8 Max
Specifications
- Input
- Output
- Context window
- 256K tokens
- Released
- Aug 2026
Performance
- Speed
- 75 t/s
- TTFT
- —
- Latency
- 200 ms
- Intelligence
- —
Pricing
- Input
- €2.50 per 1M tokens
- Output
- €6.00 per 1M tokens
€0.60 cache hit
About this model
Qwen3.8 2.4T A95B is a Mixture-of-Experts model with 2.4T total parameters and 95B activated, featuring 512 experts with 10 routed plus 1 shared per token. Built on a Gated DeltaNet architecture with 92 layers, it supports a native 262K context window extensible to over 1M tokens. The model requires thinking mode for all interactions with adjustable reasoning depth via reasoning_effort (xhigh, medium, low) and achieves 92.6 on GPQA Diamond, 86.6 on Terminal Bench 2.1, and 93.0 on PaperBench, demonstrating strong performance across coding, agentic tasks, and professional domains. Available under the Qwen license.
Technical specifications
- Capabilities
- Input modalities
- Output modalities
- Reasoning
- Default on
Knowledge horizon
Released Aug 2026
Today
Since release 0 mo
See also
Add Model to Comparison
Search for a model to add