In brief: Microsoft has released MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and the faster Flash variant. An independent test supports the transcription model's top accuracy ranking. Full voice-agent latency still includes reasoning, tool calls and several ASR capabilities that this benchmark does not cover.
Streaming across 60 languages
MAI-Transcribe-2-Streaming returns provisional hypotheses while someone speaks, then replaces them with stable segments. Microsoft lists 60 languages, continuous automatic language detection and introductory pricing of $0.54 per audio hour through year-end. The model is available in preview through Microsoft Foundry.
Artificial Analysis ranks it first for both first-partial and final-transcript accuracy. Its benchmark contains roughly eight hours of conversational, parliamentary and earnings audio, with response time measured after an automated voice-activity detector marks the end of speech. This makes comparisons more reproducible, but does not establish identical error rates on noisy phone calls, specialist terms or customer-specific accents.
Microsoft's model overview also lists important omissions: speaker diarisation, word-level timestamps and contextual keyword biasing are not supported. Contact centres may therefore need additional components or a later version.
Two output models, different latency definitions
MAI-Voice-2.1 costs $22 per million characters and lists roughly 550 milliseconds of model inference. Voice-2.1-Flash costs $15 and lists about 45 milliseconds of inference; Microsoft separately claims 150 milliseconds end-to-end for generating 45 seconds of audio. These figures cover different stages and cannot be compared directly with time to the first transcript.
Both voice models support 23 languages, according to the launch post. Microsoft's locale documentation is inconsistent: the announcement says 26 locales, while the product page visibly lists 28 regional entries. The model card describes voice cloning from five to 60 seconds of reference audio; access is gated and tied to consent safeguards.
Pandorex Analysis: a faster audio edge, not the entire agent
Microsoft reduces latency at the input and output edges and lets an application start compute from provisional text. Total conversational delay still depends on end-of-speech detection, the reasoning model, tools, networking and playback. Partial transcripts can also change, so agents should wait for stable segments or confirmation before irreversible actions.
The supported advance is an integrated audio stack with independently validated ASR accuracy. The stronger claim that it automatically produces a delay-free voice agent is not established.
