Z.ai Releases GLM-5.3-Flash MoE 320B with 1M Token Context
Z.ai launches GLM-5.3-Flash, a 320B-A18B MoE multimodal model supporting a 1M-token context window, available on Hugging Face and via API.

Stock photo for illustration only, not from the actual event
- Z.ai has released GLM-5.3-Flash, a 320B-A18B natively multimodal MoE model.
- It features a massive 1-million-token context window for extensive data processing.
- Model weights are live on Hugging Face under an MIT license alongside a hosted API.
- Scored 57 on the Artificial Analysis Intelligence Index with competitive pricing.
Z.ai has officially rolled out its latest artificial intelligence model, GLM-5.3-Flash, trained from scratch on a massive 30-trillion-token multimodal corpus. Designed to deliver high performance and cost-efficiency for enterprise and developer workloads, the model is immediately accessible through multiple tracks.
The model weights are currently hosted on Hugging Face under an open MIT license, allowing developers to download and integrate them freely. Simultaneously, a fully priced and operational hosted API is up and running to serve inference requests right out of the box.

Stock photo for illustration only, not from the actual event
Architecturally, GLM-5.3-Flash utilizes a Mixture of Experts (MoE) design combining 320 billion total parameters with 18 billion active parameters per token. Its standout capability is the native support for a 1-million-token context window, engineered to handle extensive multimodal inputs seamlessly.
Mixture of Experts (MoE) architecture allows large language models to maintain high capability while optimizing computational efficiency. By routing tokens only to specialized sub-networks, models like GLM-5.3-Flash can achieve fast inference speeds and lower operational costs compared to dense models of equivalent total parameter size, making large context windows much more practical for production environments.
Independent evaluation by Artificial Analysis awards the model a score of 57 on the Intelligence Index, delivering 48.7 output tokens per second and a TTFT of 1.52 seconds on Z.ai's API. While vision benchmarks show it trailing behind competitors like Gemini 3.7 Flash in specific tests, the intelligence-per-dollar ratio remains a primary selling point.
Standard API pricing is set at $0.15 per million input tokens, $0.03 per million cached input tokens, and $0.50 per million output tokens. The model is also integrated into all GLM Coding Plan tiers—Lite, Pro, and Max—offering triple the usable quota of GLM-5.3, with local deployment supported via SGLang, vLLM, TokenSpeed, and KTransformers.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment