三款模型均透過 Microsoft Foundry 公開預覽提供。轉錄模型每音訊小時 0.54 美元的價格屬於 2026 年底前的優惠價;這些語音元件本身並不取代語音代理所需的推理或其他服務。
What are Microsoft’s three new in-house AI models for real-time transcription and voice generation, how do their capabilities, language suppMicrosoft’s MAI lineup combines streaming speech recognition with two text-to-speech models.
AI 提示詞
Create a landscape editorial hero image for this Studio Global article: What are Microsoft’s three new in-house AI models for real-time transcription and voice generation, how do their capabilities, language supp. Article summary: Microsoft’s three models are **MAI-Transcribe-2-Streaming**, which listens and transcribes as speech arrives, and **MAI-Voice-2.1** and **MAI-Voice-2.1-Flash**, which generate spoken responses. Together, they give develo. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
openai.com
微軟這次推出的三款 MAI 模型,分別負責語音互動的「聽」與「說」:MAI-Transcribe-2-Streaming 將即時語音轉成文字;MAI-Voice-2.1 和 MAI-Voice-2.1-Flash 則把文字轉成語音。這批模型瞄準的是開發語音代理的團隊,目前均透過 Microsoft Foundry 公開預覽提供。 3639