Microsoft Releases MAI-Transcribe-2 for Live Speech
Microsoft’s new streaming model transcribes speech continuously in 60 languages, targeting voice agents that must reason before a speaker finishes.
Microsoft has released MAI-Transcribe-2-Streaming, a low-latency speech-to-text model available through Microsoft Foundry and Azure Speech. The model continuously produces partial transcripts, revises them as more audio arrives and commits a stable result when an utterance ends.
Designed for interruption-free voice systems
The model supports 60 languages with automatic language detection. Microsoft says it can produce initial hypotheses within the low hundreds of milliseconds and that words generally appear around 320 milliseconds after being spoken, compared with more than 500 milliseconds for the closest competing systems in its cited evaluation.
The distinction between partial and final transcription is important for voice agents. A system can begin classifying a request, preparing a tool call or selecting a response while the speaker is still talking. That reduces the dead time created by the traditional pipeline, in which an application waits for the full utterance before passing text to a language model.
Microsoft positions the model for call centers, voice assistants, meetings, lectures and live captions. It is offered through an OpenAI Realtime-compatible WebSocket interface as well as the Azure Speech SDK, giving developers both a lower-level integration path and managed connection handling.
Why it matters
The release is less about a dramatic new speech benchmark than about moving voice-agent latency into the application architecture. Faster partial transcripts can allow reasoning and tool preparation to overlap with speech, but early hypotheses can also be wrong and may trigger premature actions. Microsoft’s public preview status means production reliability and error behavior remain to be established. Still, the combination of multilingual coverage, streaming output and agent-oriented interfaces makes speech recognition a more active control surface for AI systems rather than a passive transcription step.