Skip to main content

Z.ai Releases GLM-5.3-Flash MoE 320B with 1M Token Context

Z.ai launches GLM-5.3-Flash, a 320B-A18B MoE multimodal model supporting a 1M-token context window, available on Hugging Face and via API.

AI-written
Inewgen
27 Aug 2026Source: MarkTechPost2 min read (0 views)
Share
Z.ai Releases GLM-5.3-Flash MoE 320B with 1M Token Context

Stock photo for illustration only, not from the actual event

Font size
  • Z.ai has released GLM-5.3-Flash, a 320B-A18B natively multimodal MoE model.
  • It features a massive 1-million-token context window for extensive data processing.
  • Model weights are live on Hugging Face under an MIT license alongside a hosted API.
  • Scored 57 on the Artificial Analysis Intelligence Index with competitive pricing.

Z.ai has officially rolled out its latest artificial intelligence model, GLM-5.3-Flash, trained from scratch on a massive 30-trillion-token multimodal corpus. Designed to deliver high performance and cost-efficiency for enterprise and developer workloads, the model is immediately accessible through multiple tracks.

The model weights are currently hosted on Hugging Face under an open MIT license, allowing developers to download and integrate them freely. Simultaneously, a fully priced and operational hosted API is up and running to serve inference requests right out of the box.

artificial intelligence data code screen

Stock photo for illustration only, not from the actual event

320BTotal Parameters
18BActive Parameters
1MMax Context Tokens

Architecturally, GLM-5.3-Flash utilizes a Mixture of Experts (MoE) design combining 320 billion total parameters with 18 billion active parameters per token. Its standout capability is the native support for a 1-million-token context window, engineered to handle extensive multimodal inputs seamlessly.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

Mixture of Experts (MoE) architecture allows large language models to maintain high capability while optimizing computational efficiency. By routing tokens only to specialized sub-networks, models like GLM-5.3-Flash can achieve fast inference speeds and lower operational costs compared to dense models of equivalent total parameter size, making large context windows much more practical for production environments.

Independent evaluation by Artificial Analysis awards the model a score of 57 on the Intelligence Index, delivering 48.7 output tokens per second and a TTFT of 1.52 seconds on Z.ai's API. While vision benchmarks show it trailing behind competitors like Gemini 3.7 Flash in specific tests, the intelligence-per-dollar ratio remains a primary selling point.

Standard API pricing is set at $0.15 per million input tokens, $0.03 per million cached input tokens, and $0.50 per million output tokens. The model is also integrated into all GLM Coding Plan tiers—Lite, Pro, and Max—offering triple the usable quota of GLM-5.3, with local deployment supported via SGLang, vLLM, TokenSpeed, and KTransformers.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article