Microsoft’s three MAI speech models pair 60 language streaming transcription—with a reported No. The transcription model returns partial results as speech arrives.
Published byEdited with GPT-6 LunaImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What are Microsoft’s three new in-house AI models for real-time transcription and voice generation, how do their capabilities, language supp. Article summary: Microsoft’s three models are **MAI-Transcribe-2-Streaming**, which listens and transcribes as speech arrives, and **MAI-Voice-2.1** and **MAI-Voice-2.1-Flash**, which generate spoken responses. Together, they give develo. Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
Microsoft’s three new MAI models cover both sides of a spoken interaction: MAI-Transcribe-2-Streaming turns live speech into text, while MAI-Voice-2.1 and MAI-Voice-2.1-Flash generate speech from text. Announced for developers building voice agents, the models are available in public preview through Microsoft Foundry. 36
39
| Model | Main role | Reported language support | Reported price |
|---|---|---|---|
| MAI-Transcribe-2-Streaming | Live speech-to-text with incremental transcript updates | 60 languages, with automatic language detection | $0.54 per audio hour, an introductory rate through the end of 2026 |
| MAI-Voice-2.1 | Expressive text-to-speech | 23 languages | $22 per million characters |
| MAI-Voice-2.1-Flash | Faster text-to-speech for responsiveness-focused applications | 23 languages reported for the voice models | $15 per million characters |
These are reported listings, not a guarantee that prices or availability will remain unchanged. In particular, Microsoft’s transcription price is described as introductory. 4
35
Unlike a transcription system that waits for someone to finish speaking, MAI-Transcribe-2-Streaming sends transcript updates while speech is still arriving. It accepts continuous audio over a WebSocket and automatically detects among 60 languages, according to reporting on its model listing. 1
The model’s latency figures vary by report and measurement. One account says it can begin producing partial results just over 100 milliseconds after receiving audio; another says words can begin appearing roughly 320 milliseconds after they are spoken. Those descriptions use different stated starting points, so they should not be read as a like-for-like latency comparison. 2
4
On the Artificial Analysis streaming-transcription benchmark, reports say the model debuted at No. 1 for accuracy. One report gives a final word-error rate of 2.5% among 38 models. That is a reported benchmark result, not a guarantee of accuracy for every language, recording condition or application. 34
The listed introductory price is $0.54 per hour of audio through the end of 2026. Developers evaluating the model should account for the fact that this is a time-limited promotional rate. 4
35
Both voice models turn text into spoken audio, but they are positioned for different priorities. MAI-Voice-2.1 is described as the more expressive option, while Flash is the faster variant for applications where keeping the wait between turns short matters. The reported language coverage is 23 languages for the voice models. 6
35
36
Reported listings price MAI-Voice-2.1 at $22 per million characters and MAI-Voice-2.1-Flash at $15 per million characters. A report also describes Flash as generating 45 seconds of audio with 150 milliseconds of latency; treat that as a reported figure rather than a universal measure of end-to-end response time. 35
36
There is no directly comparable public benchmark score for voice quality in the available reports. The transcription model’s accuracy ranking should not be taken as a ranking of either speech-generation model. 32
34
All three models are reported in public preview through Microsoft Foundry. Public preview gives developers a way to test the models, but it is not the same as a statement that they are generally available. 36
For developers, the practical change is that Microsoft now offers its own speech-recognition and speech-generation components for a voice-agent workflow. That could make it easier to build more of the speech layer within Microsoft’s developer ecosystem and reduce the need to source those particular components elsewhere. It does not establish that a complete voice agent can work without other providers: reasoning, orchestration and other capabilities may still involve different models or services. 6
39
The clearest point of comparison is the transcription model: it supports 60 languages, returns incremental results, and has a reported top Artificial Analysis accuracy ranking, alongside a stated 2.5% final word-error rate. The two voice models offer an expressive-versus-speed choice, with reported coverage of 23 languages and different per-character prices. Developers should verify current preview access, pricing and latency in their intended setup before making a production decision. 1
34
35
36
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Microsoft’s three MAI speech models pair 60 language streaming transcription—with a reported No.
Microsoft’s three MAI speech models pair 60 language streaming transcription—with a reported No. The transcription model returns partial results as speech arrives. Reports describe its latency differently: one cites just over 100 milliseconds to initial results, another roughly 320 milliseconds for words to appea...
The lineup gives developers Microsoft built speech input and output, but does not by itself replace outside models for reasoning or other parts of a voice agent.