Command Palette
Search for a command to run

GLM 5.3 Flash

by ZAI

Specifications

Input
Output
Context window
1M tokens
Released
Aug 2026

Performance

Speed
197 t/s
TTFT
Latency
133 ms
Intelligence

Pricing

Input
€0.10
per 1M tokens
Output
€0.40
per 1M tokens
€0.02 cache hit

About this model

ZAI GLM 5.3 Flash is a 320B parameter Mixture-of-Experts (MoE) language model with 18B active parameters, and the first natively multimodal model in the GLM-5 series. It introduces a hybrid architecture combining sparse and linear attention to reduce long-context serving costs, along with Manifold-Constrained Hyper-Connections (mHC) for improved scaling efficiency, trained on a 30T-token multimodal corpus. The model supports configurable reasoning effort levels (low, high, max) with thinking enabled by default, and features a context window of up to 1 million tokens. It outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. Released under the MIT license.

Technical specifications

Capabilities
Input modalities
Output modalities
Reasoning
Default on

Knowledge horizon

Released Aug 2026
Today
Since release 0 mo

See also

Add Model to Comparison
Search for a model to add
Command Palette
Search for a command to run