ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
MAI-Transcribe-2-Streaming launches: Microsoft's voice model writes while it listens

MAI-Transcribe-2-Streaming launches: Microsoft's voice model writes while it listens

AI information • Admin • • 5 views

MAI-Transcribe-2-Streaming was released on October 1, 2026 by Microsoft AI through its official blog, alongside the speech synthesis models MAI-Voice-2.1 and the faster MAI-Voice-2.1-Flash. Microsoft packages the three as a voice-agent kit: one listens, one speaks, and the time left for the reasoning model in between is deliberately squeezed.

Writing while listening, acting before the sentence ends

MAI-Transcribe-2-Streaming is Microsoft's first streaming transcription model. Instead of waiting for a full sentence, it produces its first partial transcripts in just over 100 milliseconds from a continuous audio stream, revises them as more context arrives, and locks in the final text the moment an utterance ends. In Microsoft's dictation and subtitling evaluations, words appear about twice as fast as with its closest competitor, and on Artificial Analysis it ranks first for accuracy on both final and partial transcripts. It supports 60 languages with automatic, continuous language detection, at an introductory price of $0.54 per hour of audio through the end of 2026.

One voice, speaking 23 languages

MAI-Voice-2.1 is Microsoft's strongest multilingual text-to-speech model so far, supporting 23 languages and 26 locales at $22 per million characters. Its signature trick is keeping one voice identity across languages: the same speaker in English, Mandarin and German still sounds like the same person, with a native accent in each. The Flash variant supports the same languages, generates 45 seconds of audio with about 150 milliseconds of end-to-end latency, runs inference 55% faster, and costs $15 per million characters. Both voice models can clone a voice from a few seconds of reference audio, with built-in consent guardrails against misuse.

Why Microsoft built its own ears and mouth

A voice agent is a loop: hear, understand, decide, speak, all inside the rhythm of a natural conversation. Shaving a few hundred milliseconds off both listening and speaking buys time for reasoning and tool calls in the middle. Customer-service agents can start identifying a request before the caller finishes, multilingual assistants can detect the language and reply in the same voice, and tutoring products no longer need a new teacher for each language. One caveat: the 100-millisecond figure is the model's time to a partial transcript, not the latency of the whole voice pipeline, which networks and applications still affect.

Recommended Tools

More