Press Esc to close

Microsoft AI's first streaming transcription model tops accuracy leaderboard, adds two voices

Official Microsoft AI artwork of pink comma shapes swirling across a blurred blue background with soft streaks of light, like sound moving through the air

Image: Microsoft

Microsoft AI has released its first streaming transcription model, MAI-Transcribe-2-Streaming, and it claims to be the most accurate real-time speech-to-text model on Artificial Analysis. Two new voice models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, launched alongside it.

Instead of waiting for you to finish talking, the model starts guessing what you said (Microsoft calls these guesses "partials") just over 100ms after it hears audio, then corrects them as more context comes in. That lets a voice agent start reasoning or calling tools mid-sentence. It works in 60 languages and detects which one you're speaking automatically.

Microsoft says it ranks first on Artificial Analysis for both final and partial transcripts. Its chart, based on the Artificial Analysis streaming leaderboard, puts the model's final word error rate at 2.5%, ahead of Grok Voice Transcribe 2.0 (2.73%) and Muse Voice Transcribe (3.06%). Microsoft's internal tests also show words appearing twice as fast as with its closest competitor. Pricing starts at an introductory $0.54 per hour of audio through the end of the year.

Microsoft bar chart of the Artificial Analysis streaming word error rate leaderboard. MAI-Transcribe-2-Streaming has the lowest final error rate at 2.5 percent, ahead of Grok Voice Transcribe 2.0, Muse Voice Transcribe and 25 other streaming models, with a dotted line marking each model's first partial error rate

Image: Microsoft (data: Artificial Analysis)

Two new voices, too

MAI-Voice-2.1 handles the talking. It supports 23 languages and 26 locales, and a single voice can switch between them with a native accent, so a tutoring app could jump from English to Mandarin without changing who's speaking. It costs $22 per 1 million characters.

MAI-Voice-2.1-Flash is the faster, cheaper option for high-volume work. Microsoft says it can generate 45 seconds of audio with just 150ms of end-to-end latency, and that it runs inference 55% faster and costs about 60% less than comparable models, at $15 per 1 million characters.

Both voice models can clone a voice from a few seconds of reference audio, with consent guardrails built in. In a 4,000-listener Turing test, 50.3% rated MAI-Voice as equally or more human-like than real human recordings.

Pair the fast transcriber with Flash, Microsoft argues, and a voice agent spends less time waiting and more time thinking. The new voices follow MAI-Voice-2-Flash, which we covered in July.

All three models are available through Microsoft Foundry, the MAI Playground (which has a new demo agent called Chatter), Vercel and Azure Voice Live, with LiveKit coming soon. The two voice models are also on OpenRouter.

Comments