Advertisement 780 x 90
Microsoft’s new AI can transcribe speech before the speaker finishes
AI Generated

Microsoft’s new AI can transcribe speech before the speaker finishes

Microsoft’s new AI can transcribe speech before the speaker finishes
Microsoft’s MAI-Transcribe-2-Streaming sends transcript updates while people are still talking. Here’s what the preview offers, who can use it and why the early results could make voice apps feel faster.

Microsoft’s streaming transcription model

Microsoft’s streaming transcription model

Microsoft is aiming to make voice assistants feel less like turn-taking machines. Its new MAI-Transcribe-2-Streaming model sends text back while a person is still speaking, rather than waiting for a pause or the end of a recording.

That small change could matter for live captions, dictation and voice assistants. The model can return an initial transcript hypothesis just over 100 milliseconds after receiving audio, then revise it as more speech arrives. Microsoft says developers can use those early results to let a voice agent begin reasoning or calling a tool before the speaker finishes.

Announced on October 1, the model is available in public preview through Microsoft Foundry. It supports 60 languages and automatic language detection, according to Microsoft. The company says it ranks first for accuracy on Artificial Analysis’ evaluations of both final and partial transcripts. Those rankings and performance claims come from Microsoft’s announcement; they are not a guarantee of results in every real-world setting.

MAI-Transcribe-2-Streaming brings live text to voice apps

Traditional transcription systems commonly return a finished transcript after processing a chunk of audio. Streaming transcription instead sends successive updates as the audio arrives. Early words may change as the system gets more context, so applications need to handle provisional text as well as stable results.

That approach could make a voice assistant more responsive. A system might start preparing an answer while a user is still finishing a request. Live captions could appear sooner, and speech-to-text tools could show words as someone dictates. Microsoft also points to meeting and lecture captioning, call centers and voice-driven interfaces as possible uses.

The model is aimed at developers, not a new consumer app that people can simply download. Microsoft documents access through Foundry, including a real-time API and an Azure Speech SDK integration. The documentation labels the offering a public preview and cautions that preview services do not come with a service-level agreement or a recommendation for production workloads.

Three models cover both sides of a conversation

Microsoft paired the transcription release with MAI-Voice-2.1 and MAI-Voice-2.1-Flash, two text-to-speech models. The standard version supports 23 languages and 26 locales, according to the company. Microsoft says one voice can keep a consistent identity while speaking different languages, a feature that could be useful for multilingual lessons or assistants.

The Flash version is designed for high-volume, latency-sensitive speech generation. Microsoft says it can generate 45 seconds of audio with end-to-end latency of 150 milliseconds. Those figures are company-reported performance claims, not an independent comparison.

The combination points toward a broader product goal: letting developers build voice agents that can listen and respond with less delay. Transcription handles incoming speech; text-to-speech produces the spoken reply. In practice, the experience will depend on the full application, including network conditions, how it handles interruptions and whether its answers are useful.

Preview pricing, with a developer-focused caveat

Microsoft lists an introductory price of $0.54 per hour of audio for MAI-Transcribe-2-Streaming through the end of 2026. The company prices MAI-Voice-2.1 at $22 per million characters, and the Flash model at $15 per million characters. Actual costs depend on usage and how a developer builds and deploys an application.

The central takeaway is not that speech recognition has suddenly become flawless. It is that Microsoft is making early transcript updates a building block for more immediate voice interactions. Developers can explore the models in preview now, but the next test is whether the speed and accuracy hold up in everyday products, across languages and noisy environments.

Source: Microsoft AI — microsoft.ai/news/our-first-streaming-transcription-model

Advertisement 780 x 90

What do you think?

+0 Points

What's your reaction?

0
AWESOME!
0
NICE
0
LOVED
0
LOL
0
FUNNY
0
FAIL!
0
OMG!
0
EW!

Comments

G

0 comment