GLM 5.3 Flash
Specifications
- Input
- Output
- Context window
- 1M tokens
- Veröffentlicht
- Aug 2026
Performance
- Speed
- 197 t/s
- TTFT
- —
- Latency
- 133 ms
- Intelligence
- —
Pricing
- Eingabe
- €0.10 per 1M tokens
- Ausgabe
- €0.40 per 1M tokens
Über dieses Modell
ZAI GLM 5.3 Flash is a 320B parameter Mixture-of-Experts (MoE) language model with 18B active parameters, and the first natively multimodal model in the GLM-5 series. It introduces a hybrid architecture combining sparse and linear attention to reduce long-context serving costs, along with Manifold-Constrained Hyper-Connections (mHC) for improved scaling efficiency, trained on a 30T-token multimodal corpus. The model supports configurable reasoning effort levels (low, high, max) with thinking enabled by default, and features a context window of up to 1 million tokens. It outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. Released under the MIT license.
Technische Daten
- Fähigkeiten
- Eingabe-Modalitäten
- Ausgabe-Modalitäten
- Reasoning
- Standard on