Alibaba Qwen Releases Qwen3.8-Omni-Flash Omni-Modal Model
Alibaba introduces Qwen3.8-Omni-Flash, a 1M-context omni-modal model featuring native audio-video understanding and agentic capabilities.

Stock photo for illustration only, not from the actual event
- Alibaba launches Qwen3.8-Omni-Flash supporting a 1M token context window
- Agent-driven video and audio processing focuses compute on relevant segments
- Boosts OmniVideoBench accuracy to 67.8 while cutting token usage by 45.7%
- Available now as a hosted API on QwenCloud and Alibaba Cloud Model Studio
The Alibaba Qwen research team has officially released Qwen3.8-Omni-Flash, the company's first omni-modal model built specifically around agentic capabilities. The model combines native audio-video understanding, advanced reasoning, and tool usage into a single unified architecture.
Regarding deployment, the model is currently available as a hosted API on QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio. It is built upon the Qwen3.8-Flash-Next base architecture. However, no open weights were announced at launch, meaning self-hosting is not an option at this time.
A major highlight is its 1-million-token context window. QwenCloud documentation lists a maximum input of 991K tokens and a maximum output of 131K tokens, with a maximum reasoning length of 262K tokens. Output is strictly text-only; developers requiring generated speech are directed to use Qwen3.5-Omni instead. Thinking is enabled by default with the reasoning effort set to xhigh.

Stock photo for illustration only, not from the actual event
In terms of video processing architecture, conventional video models read long files sequentially from start to finish. In contrast, the Qwen research team adopted an agent-driven approach where the agent starts from the prompt, decides what to watch and hear, and gathers evidence across several coarse-to-fine rounds, directing compute and tokens strictly to relevant segments.
Benchmark results on OmniVideoBench show accuracy rising from 63.4 to 67.8, while token consumption drops from 145,736 to 79,117—representing an approximately 45.7% reduction in tokens. Furthermore, audio-visual performance closely approaches Gemini 3.8 Flash, with overall audio performance claimed to surpass it.
"Meet Qwen3.8-Omni-Flash, Qwen's first omni-modal model built around agentic capabilities! Native audio-video understanding, reasoning, and tool use come together in one model"
Qwen Research Team
On the pricing side, QwenCloud lists $0.15 per 1M input tokens and $0.47 per 1M output tokens, with implicit cache hits costing $0.016 per 1M tokens. This delivers substantial cost reductions compared to Qwen3.5-Omni-Plus, lowering audio input costs by over 98% per hour and audio-visual input costs by over 93% per hour.
The introduction of Qwen3.8-Omni-Flash highlights Alibaba's strategic shift toward agentic multimodal workflows. By utilizing selective attention mechanisms rather than brute-force sequential processing for long-form video, the model drastically cuts operational costs and token overhead, setting a new standard for efficient enterprise-grade multimodal AI deployment.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment