Skip to main content

Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model

Alibaba launches Qwen3.8-Max, a 2.4 trillion parameter Mixture of Experts model featuring a 1M-token context window, flexible API pricing, and a smaller 27B variant for local hardware.

AI-written
Inewgen
03 Aug 2026Source: MarkTechPost3 min read (0 views)Last updated 29 Aug 2026
Share
Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model

Stock photo for illustration only, not from the actual event

Font size
  • Alibaba announces Qwen3.8-Max, a 2.4 trillion parameter MoE model
  • Supports up to a 1M-token context window and 131K maximum output tokens
  • Includes the Qwen3.8-27B checkpoint designed for ordinary on-premise GPUs
  • API pricing starts at $2.00 per 1M input tokens and $6.00 per 1M output tokens

Alibaba has officially introduced its newest artificial intelligence model, Qwen3.8-Max, marking the most capable release in the Qwen family to date. Built on a Mixture of Experts (MoE) architecture, the model boasts a staggering total of 2.4 trillion parameters, engineered to handle complex reasoning tasks and deep contextual understanding at scale.

Deployment options vary depending on the specific artifact being applied. The hosted API is readily deployable for organizations of any size, offering OpenAI- and DashScope-compatible integration where switching requires only a base-URL and model-ID change. Conversely, the open-weight checkpoint sits at 2.4 trillion parameters, functioning as a multi-node datacenter asset. Because Alibaba has not disclosed the activated-parameter count, exact serving costs cannot yet be modeled. To bridge this gap for local infrastructure, the Qwen3.8-27B checkpoint is provided to fit standard on-premise GPU hardware.

advanced artificial intelligence futuristic server rack

Stock photo for illustration only, not from the actual event

The published feature set directly addresses four key industries: software engineering, legal and financial document review, media and e-commerce operations, and design. Intended applications range from repository-scale coding agents and long-document knowledge bases to long-video indexing, structured data extraction, and multi-step research assistants.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

2.4TTotal MoE Parameters
1MToken Context Window
131KMaximum Output Tokens

Technical specifications outline a 1M-token context window capability, with a maximum input of 991K tokens that drops slightly to 983K when thinking is enabled. Maximum output reaches 131K tokens across both modes, backed by a reasoning budget of up to 262K tokens. Rate limits are capped at 2 million tokens per minute and 15,000 requests per minute.

The rollout of Qwen3.8-Max highlights the ongoing industry push toward ultra-large MoE architectures that balance massive parameter scales with optimized inference economics. By offering both cloud-hosted APIs and a smaller 27B local variant, Alibaba is targeting a broad spectrum of deployment needs—from enterprise-grade cloud integration to specialized on-premise execution—while leveraging advanced caching mechanics to manage operational expenses efficiently.

Pricing for the API is structured at $2.00 per 1 million input tokens and $6.00 per 1 million output tokens. Implicit cache reads cost $0.25 per 1 million tokens, while explicit cache creation and reads are priced at $2.50 and $0.17 per 1 million tokens, respectively. Because cached input runs eight times cheaper than fresh input, prefix stability plays a more significant role in driving overall costs than raw prompt length. Supported features include function calling, structured outputs, batches, prefix completion, and fine-tuning, alongside five built-in tools on the Responses API: code_interpreter, web_search, web_extractor, t2i_search, and i2i_search.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article