Skip to main content

Alibaba Releases Qwen3.8-LiveTranslate Cutting Lag to 2.3s

Alibaba Qwen Team launches Qwen3.8-LiveTranslate, a real-time interpretation model supporting 60 languages with an average lag cut to 2.3 seconds via API.

AI-written
Inewgen
Live20 Sep 20262 min read (0 views)
Share
Alibaba Releases Qwen3.8-LiveTranslate Cutting Lag to 2.3s

Stock photo for illustration only, not from the actual event

Font size
  • Alibaba launches Qwen3.8-LiveTranslate, a real-time simultaneous interpretation model.
  • Cuts average LAAL latency down to 2.3 seconds via WebSocket API integration.
  • Supports 60 source languages with speech output capability in 29 languages.

The Alibaba Qwen team has officially released Qwen3.8-LiveTranslate, a state-of-the-art real-time interpretation model built to address the classic tradeoff in simultaneous translation between waiting for complete context and minimizing listener delay. The model rebuilds this processing loop utilizing an innovative Interleave architecture.

Building on the Qwen-Omni stack, this new release integrates large-scale multimodal data, cross-language and cross-modal alignment, alongside visual enhancement features. This allows the system to process both audio inputs and optional visual cues—such as speaker gestures, lip movements, and on-screen text—to maintain high accuracy in noisy environments or when encountering ambiguous terminology, with documentation recommending no more than two images per second.

artificial intelligence cloud computing server rack

Stock photo for illustration only, not from the actual event

2.3sAverage LAAL Latency
60Supported Languages
29Speech Output Languages

Regarding language coverage, the model comprehends 60 languages in total, returning both audio and text for 29 of them while outputting text-only for the remaining 31. Supported speech output includes Chinese, English, Arabic, German, French, Spanish, Japanese, Korean, and Hindi. Furthermore, development teams can configure up to 1,000 hotwords to map specific source terminology directly to fixed target translations.

Implementing an Interleave architecture within a real-time multimodal framework represents a major step forward in overcoming traditional translation bottlenecks. By fusing audio and visual streams concurrently, the model significantly enhances disambiguation capabilities in complex acoustic environments without accumulating excessive latency.

Developers can integrate the model via the WebSocket Realtime API using the model ID qwen3.8-livetranslate-flash-realtime, utilizing default speaker detection and the default Tina voice profile. The system operates with a 53,248-token context window and standard rate limits of 10 requests and 100,000 tokens per minute.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article