三款模型均通过 Microsoft Foundry 公共预览提供。它们补齐了语音输入与输出组件,但不代表语音代理的推理等其他部分也无需外部模型。
What are Microsoft’s three new in-house AI models for real-time transcription and voice generation, how do their capabilities, language suppMicrosoft’s MAI lineup combines streaming speech recognition with two text-to-speech models.
AI 提示
Create a landscape editorial hero image for this Studio Global article: What are Microsoft’s three new in-house AI models for real-time transcription and voice generation, how do their capabilities, language supp. Article summary: Microsoft’s three models are **MAI-Transcribe-2-Streaming**, which listens and transcribes as speech arrives, and **MAI-Voice-2.1** and **MAI-Voice-2.1-Flash**, which generate spoken responses. Together, they give develo. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
openai.com
微软推出三款面向语音交互的 MAI 模型:MAI-Transcribe-2-Streaming 将实时语音转成文字,MAI-Voice-2.1 和 MAI-Voice-2.1-Flash 则把文字转换为语音。它们面向希望构建语音代理的开发者,目前均通过 Microsoft Foundry 公共预览提供。 3639