Qwen has released Qwen3.8-LiveTranslate, a new real-time simultaneous interpretation model designed to improve translation quality while reducing the delay between speech and translated output. The model uses a new Interleave architecture and reduces average lagging, or LAAL, from 2.8 seconds in the previous generation to 2.3 seconds, according to Qwen.
The September 18, 2026 release also adds three capabilities aimed at more complex conversations: real-time speaker separation, synchronized source-and-translation output, and long-context disambiguation. Qwen says the model supports 60 languages for audio input and text output, while speech output is available in 29 languages.
Rather than treating simultaneous interpretation as a simple speech-recognition and translation pipeline, Qwen says Qwen3.8-LiveTranslate combines audio and text into an interleaved stream so that previously heard audio and generated translations can be reused during the conversation.
What Qwen3.8-LiveTranslate Changes?
Real-time interpretation requires a balance between translation quality and latency. A system that waits too long to understand an entire sentence may produce better context but creates a noticeable delay, while translating too aggressively can result in incomplete or less coherent output.
Qwen3.8-LiveTranslate is designed around this trade-off. Qwen says its Interleave architecture improves faithfulness, fluency and conciseness while reducing average lagging from 2.8 seconds to 2.3 seconds.
The model uses a Hybrid-MoE-based Thinker–Talker architecture. The Thinker processes video, audio, source text and translation as a temporally ordered causal sequence, while the Talker uses the translation and source audio to synthesize speech that can preserve characteristics of the original speaker’s voice.
This architecture is particularly relevant to simultaneous interpretation because the model has to understand incoming speech, generate a translation and produce spoken output continuously rather than waiting for a complete recording.
Real-Time Speaker Separation and Voice Preservation
One of the major additions in Qwen3.8-LiveTranslate is real-time speaker separation.
In a multi-person conversation, the model can distinguish between speakers and associate individual sentences with the person who said them. Qwen also says the system can preserve each speaker’s timbre more consistently when generating translated speech.
This could be particularly useful for meetings, interviews and other multi-speaker scenarios where a conventional translation system may produce a continuous stream without making speaker attribution clear.
Qwen’s official evaluation includes a multi-speaker long-audio test covering 14 language directions. The company reports that Qwen3.8-LiveTranslate performs ahead of the mainstream real-time interpretation systems it compared against across translation faithfulness, fluency, conciseness and diarization error rate.
These are Qwen-reported benchmark results, so they should be understood within the datasets and evaluation methodology selected by the company.
Source and Translation Can Appear Together
Qwen3.8-LiveTranslate also introduces synchronized source-and-translation output, allowing the original speech content and translated version to appear together on a bilingual display.
The feature can make real-time interpretation easier to verify because users do not have to rely exclusively on the translated output. Seeing the source and translation together can also support applications such as live subtitles, meeting records and multilingual video workflows.
Qwen describes synchronized bilingual output as a foundation for additional applications involving subtitle display, content organization and retrieval.
Long-Context Disambiguation Improves Names and Terminology
Another new capability is long-context disambiguation.
The system uses information from previous turns of a conversation to interpret the current speech. This is intended to reduce ambiguity around proper names, references and specialized terminology that may be difficult to translate accurately when each sentence is processed in isolation.
For example, a person’s name or technical term that has already been established earlier in a conversation can provide context for later references. Maintaining that information across turns can also help the translation remain more consistent throughout a longer discussion.
This contextual approach is different from simply increasing the number of languages supported because it addresses ambiguity within the conversation itself.
Qwen3.8-LiveTranslate Supports 60 Languages
Qwen says the model supports 60 languages for audio input and text output. The supported list includes English, Chinese, Hindi, Japanese, Korean, French, German, Spanish, Portuguese, Arabic, Bengali, Gujarati, Marathi, Punjabi, Urdu and several other languages.
Speech output is supported in 29 languages, including English, Chinese, German, Italian, Portuguese, Spanish, Japanese, Korean, French, Russian, Hindi and others.
Qwen also evaluated real-time multilingual performance across 70 language directions using the public FLEURS audio test set. The company reports improvements over its previous generation and other systems it evaluated in translation quality, average lagging, speech recognition accuracy and speech synthesis quality.
Qwen3.8-LiveTranslate API and Availability
The model is available through Alibaba Cloud’s Model Studio as qwen3.8-livetranslate-flash-realtime. The model accepts audio and image inputs and produces text and audio outputs through a WebSocket Realtime API.
The current Alibaba Cloud documentation lists a 49,152-token maximum input length, a 4,096-token maximum output length and a 53,248-token context window. It also lists the model as supporting 60 languages.
The international Singapore pricing currently listed by Alibaba Cloud is based on token usage: $7.50 per million audio-input tokens, $0.55 per million image-input tokens, $20 per million text-output tokens and $30 per million audio-output tokens.
The model does not currently support function calling or web search, according to the Model Studio documentation. That distinction matters for developers considering Qwen3.8-LiveTranslate specifically for interpretation rather than broader autonomous voice-agent workflows.
Where Qwen3.8-LiveTranslate Fits?
Qwen3.8-LiveTranslate is positioned primarily as a real-time speech and audiovisual translation system, rather than a general-purpose conversational model. Its combination of lower latency, speaker separation, bilingual display and contextual disambiguation is designed for situations where translation has to happen continuously while preserving information about who is speaking and what has already been discussed.
For developers, the release also expands Qwen’s real-time translation model family. Alibaba Cloud’s current documentation lists Qwen3.8-LiveTranslate alongside the earlier Qwen3.5-LiveTranslate models, with the new version maintaining 60-language coverage while adding the new interpretation features.
The immediate technical change is therefore not simply a larger language list. Qwen3.8-LiveTranslate combines lower reported latency with more context-aware and speaker-aware interpretation, giving developers a model designed for live multilingual conversations rather than offline translation alone.
Also Read –
Qwen3-Coder-Next: Agent-Centric Coding Model for Developers
Qwen3-TTS: Open-Source Multilingual Text-to-Speech
Qwen-Image-2512: Strongest Open-Source AI Image Model


