Alibaba highlighted Qwen3.8 LiveTranslate and five Qwen Audio 3.1 models at Apsara 2026. The Audio 3.1 lineup covers transcription, text to speech, live voice conversations, richer audio understanding and script based soundscape creation.
Published byEdited with GPT-6 SolImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What audio models did Alibaba unveil at the 2026 Apsara Conference, and what are their capabilities and intended uses—including Qwen3.8-Live. Article summary: At the 2026 Apsara Conference, Alibaba presented Qwen3.8-LiveTranslate for simultaneous interpretation and its Qwen-Audio-3.1 lineup for speech recognition, synthesis, live interaction and audio creation. The reported tr. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
Alibaba’s audio announcements around the 2026 Apsara Conference address two different jobs: translating a speaker while they are still talking, and building applications that recognize, generate, converse through or create audio. Qwen3.8-LiveTranslate handles the first; the five-model Qwen-Audio-3.1 lineup covers the second. 3
14
Qwen3.8-LiveTranslate is designed for simultaneous interpretation, producing translated output as speech arrives. Alibaba says its Interleave architecture improves translation fidelity, fluency and brevity while reducing LAAL, a measure of average translation lag, from about 2.8 to 2.3 seconds—a half-second reduction, or roughly 18%. Those quality and speed improvements are company-reported claims, not independent test results. 3
7
Alibaba also describes features for conversations involving more than one person: real-time speaker separation, a synchronized display of source and translated text, and use of earlier context to resolve ambiguous language. Its Model Studio documentation lists simultaneous interpretation and live translation as uses for the real-time model. 7
1
Qwen describes Audio 3.1 as three upgraded capabilities—ASR, TTS and Realtime—plus two new models, ASR-Next and TTS-Next. They serve distinct purposes: 14
TTS-Next is the clearest distinction between ordinary speech synthesis and audio production in this release. Alibaba pitches its script-to-soundscape capability for audiobooks, film and television, podcasts and games. It is an intended creative use, not evidence that a generated scene will meet every production requirement without editing. 3
24
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Alibaba highlighted Qwen3.8 LiveTranslate and five Qwen Audio 3.1 models at Apsara 2026.
Alibaba highlighted Qwen3.8 LiveTranslate and five Qwen Audio 3.1 models at Apsara 2026. The Audio 3.1 lineup covers transcription, text to speech, live voice conversations, richer audio understanding and script based soundscape creation.