Microsoft Launches MAI-Transcribe-2-Streaming and Two More Speech Models


microsoft Microsoft MAI Transcribe 2 Streaming
Image credit: Microsoft

Microsoft MAI speech models are expanding with three new releases aimed at real-time transcription, speech generation, and low-latency conversational voice agents.

Microsoft recently introduced MAI-Transcribe-2, and it is now following that release with MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash.

Microsoft adds real-time transcription with MAI-Transcribe-2-Streaming

The biggest change comes with MAI-Transcribe-2-Streaming, Microsoft’s new real-time speech recognition model.

Unlike MAI-Transcribe-2, which processes recorded audio, the streaming model continuously updates the transcript while a person speaks.

Microsoft says the model can return partial transcription results just over 100ms after receiving audio, making it suitable for conversational agents and other applications where responses need to appear almost immediately.

MAI-Transcribe-2-Streaming supports 60 languages and continuously detects which language someone is speaking without requiring developers to specify it beforehand.

Microsoft also says the model ranks first on Artificial Analysis for both partial and final transcription accuracy.

The company has priced MAI-Transcribe-2-Streaming at $0.54 per hour through the end of 2026.

MAI-Voice-2.1 supports 23 languages and voice cloning

Microsoft also introduced MAI-Voice-2.1, an updated speech generation model designed for multilingual voice applications.

The model supports 23 languages across 26 locales and can switch between languages while maintaining the same speaker identity.

MAI-Voice-2.1 also supports voice cloning using only a few seconds of reference audio, allowing applications to reproduce a speaker’s voice across supported languages.

Microsoft charges $22 per one million characters generated with MAI-Voice-2.1.

MAI-Voice-2.1-Flash cuts latency and costs

For workloads that prioritize speed and scale, Microsoft has released MAI-Voice-2.1-Flash.

The company says the Flash model can generate 45 seconds of audio with around 150ms of end-to-end latency.

Microsoft claims MAI-Voice-2.1-Flash delivers 55% faster inference while costing around 60% less than comparable models.

The model costs $15 per one million characters and includes the same voice cloning capabilities available with MAI-Voice-2.1.

Combined with MAI-Transcribe-2-Streaming, the Voice models give developers Microsoft’s own speech recognition and speech generation stack for building real-time conversational agents.

The new MAI speech models are already available

Microsoft has made all three models available through Microsoft Foundry, MAI Playground, and Vercel.

MAI-Voice-2.1 and MAI-Voice-2.1-Flash are also available through OpenRouter, while LiveKit support is coming soon.

Developers can also use the models through Azure Voice Live, giving Microsoft another way to integrate its MAI speech technology into real-time voice applications.

The releases expand Microsoft’s growing MAI ecosystem. The company recently published the first draft of its MAI model code of conduct and has also been expanding native voice agents in Microsoft Foundry.

More about the topics: AI, microsoft

Readers help support Windows Report. We may get a commission if you buy through our links. Tooltip Icon

Read our disclosure page to find out how can you help Windows Report sustain the editorial team. Read more

User forum

0 messages