Model release
Z.ai releases GLM-5.3-Flash, its first natively multimodal GLM-5 model
Z.ai released GLM-5.3-Flash on August 26, 2026. It is a mixture-of-experts model with 320 billion total parameters and 18 billion activated per token, on a hybrid architecture combining linear and sparse attention, with a 1M token context window. It accepts video, image, text, and file input, and Z.ai highlights native visual capability for observing interfaces and interaction feedback. On the GLM Coding Plan it carries three times the quota of GLM-5.3.
- Why it matters
- It is the first model in the GLM-5 line to take image and video input natively rather than through a separate vision variant, and the interface-observation framing aims it squarely at computer-use and browser agents. The larger coding-plan quota makes it materially cheaper than GLM-5.3 for sustained agent runs.
- Who should care
- Developers building computer-use or browser agents, and teams evaluating open-weight multimodal models
- What you can do
- Available through the Z.ai API and the GLM Coding Plan. Calls made off-peak, including all day at weekends, consume half the standard points.

