Alibaba Releases Qwen-Audio-3.1-Realtime Voice Model
Alibaba launches Qwen-Audio-3.1-Realtime, a full-duplex voice model with 262K context and companion ASR-Flash-Filetrans on September 28, 2026.

Stock photo for illustration only, not from the actual event
- Alibaba launches Qwen-Audio-3.1-Realtime, a full-duplex voice model available via managed API on QwenCloud
- Supports a 262K context window with function calling and web search capabilities
- Introduces Qwen-Audio-3.1-ASR-Flash-Filetrans for offline long-audio transcription
- Achieves improved performance metrics on Full-Duplex-Bench v3.0 with reduced filler rates
Alibaba has expanded its artificial intelligence lineup with the introduction of Qwen-Audio-3.1-Realtime, a full-duplex voice model trained to think, act, and decide when to speak. The system is currently live as a managed API via WebSocket on QwenCloud, though no open weights were announced upon its release.
The model accommodates both text and audio for inputs and outputs, boasting a context capacity of 262,000 tokens consisting of a 245,000 max input and a 16,000 max output. Default operational limits are established at 60 requests and 100,000 tokens per minute. Architecturally, the system runs two models utilizing identical Audio Encoder and LLM designs to handle full-duplex decisions, speech-to-text conversion, and streaming voice rendering.

Stock photo for illustration only, not from the actual event
Training for the model is structured across three distinct layers: Think, Act, and Speak and Coordinate. The process utilizes Core-Cocktail SFT to re-anchor the audio model to its source text LLM using million-hour-scale paired data, followed by Multimodality OPD and domain expert training via GRPO algorithms before merging into a unified deployable model.
Alongside the flagship release, the company unveiled a companion model named Qwen-Audio-3.1-ASR-Flash-Filetrans, which targets offline long-audio transcription tasks. It supports hot words, speaker separation, punctuation, and multilingual as well as Chinese dialect recognition, priced at 0.15 dollars for input and 0.47 dollars for output per 1 million tokens.
Developing full-duplex voice models represents a critical evolution in conversational AI, shifting the paradigm from turn-based exchanges to continuous, overlapping human-like interaction. By integrating a dedicated decision model that controls turn-taking dynamically, Alibaba addresses one of the most persistent bottlenecks in real-time spoken dialogue systems.
Benchmark evaluations show trade-offs in performance. On Full-Duplex-Bench v1.5, replies directed at third parties dropped from 0.13 to 0.03, while filler rates on v3.0 decreased from 0.7590 to 0.2960. However, post-interruption unwanted resumes rose from 0.035 to 0.130, and interruption stop latency clocked in at 1.116 seconds compared to 0.383 for GPT-Realtime-2.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment