

GLM-5.3 Flash
#34 in Frontier-SpraakmodelleZhipu · v5.3 · flash · 15× · tolest 29. Aug. 2026
GLM-5.3-Flash is a Mixture-of-Experts language model released by Z.ai (Zhipu AI) on August 26, 2026, featuring 320 billion total parameters and 18 billion active parameters per Token. It is the first natively multimodal model in the GLM-5 series, supporting text, image, and video inputs, and features a 1-million-Token context window with a hybrid sparse-plus-linear attention architecture. The model was previously tested anonymously as "Ox Alpha" on OpenRouter and OpenCode and is distributed under MIT license with open weights on Hugging Face. Z.ai positions it as a cost-effective alternative for coding, agent, and visual tasks that, according to the manufacturer, approaches Claude Opus 4.8 performance.
Features
| Key Benchmark (%) | Terminal-Bench 2.1: 84.3% (Claude Opus 4.8: 85.0%); DeepSWE v1.1: 63.4% |
| Context Window (Tokens) | 1,048,576 tokens (approx. 1M), max 131,072 output tokens |
| License | MIT license (open weights) |
| Multimodality | Natively multimodal: text, image and video input, text output |
| Platform | Hugging Face (zai-org/GLM-5.3-Flash), Z.ai API, OpenRouter, GLM Coding Plan |
| Price per 1M Tokens | $0.15 input / $0.50 output (launch promo through Sept 9, 2026: $0.075/$0.25); cache read $0.03 |
| Release Date | August 26, 2026 |