

GLM-5.3-Flash
#1 in Multimodal ModelsZ Ai · v5.3 · flash · 32× · last seen Aug 29, 2026
100
Momentum
GLM-5.3-Flash is a natively multimodal Mixture-of-Experts model from the GLM-5 series released by Z.ai (Zhipu AI) on August 26, 2026, with 320 billion total parameters and 18 billion active parameters per Token. It uses a hybrid sparse and linear attention architecture, supports a context window of up to 1,048,576 Token, and processes text, images, and video as input. The model weights are available under MIT license on Hugging Face; according to Z.ai, the model outperforms its predecessor GLM-5.2 at approximately one-tenth of the cost and approaches the performance of Claude Opus 4.8 on coding and agent Benchmarks.
Momentum trend
01.06.30.08.
Features
| Key Benchmark (%) | Artificial Analysis Intelligence Index: 57; Terminal-Bench 2.1: 84.3; DeepSWE v1.1: 63.4 |
| Context Window (Tokens) | 1,048,576 tokens (max output: 131,072 tokens) |
| License | MIT license (open weights) |
| Multimodality | Natively multimodal: text, image and video input, text output |
| Platform | Z.ai API Platform, Hugging Face, OpenRouter; self-hosting via SGLang, vLLM, TokenSpeed, KTransformers |
| Price per 1M Tokens | $0.15 input / $0.50 output (Z.ai list price); launch discount until Sept 9, 2026: $0.075 / $0.25 |
| Release Date | August 26, 2026 |