Pandorex
AI & Chips

Microsoft Pairs Streaming Transcription With New Voices for AI Agents

Published Pandorex Redaktion·2 min read
—
Illustration: speech flows from a microphone through provisional and stable transcripts into a violet AI processor and through a guarded voice gate to a speaker.
Editorial illustration · Pandorex

In brief: Microsoft has released MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and the faster Flash variant. An independent test supports the transcription model's top accuracy ranking. Full voice-agent latency still includes reasoning, tool calls and several ASR capabilities that this benchmark does not cover.

Streaming across 60 languages

MAI-Transcribe-2-Streaming returns provisional hypotheses while someone speaks, then replaces them with stable segments. Microsoft lists 60 languages, continuous automatic language detection and introductory pricing of $0.54 per audio hour through year-end. The model is available in preview through Microsoft Foundry.

Artificial Analysis ranks it first for both first-partial and final-transcript accuracy. Its benchmark contains roughly eight hours of conversational, parliamentary and earnings audio, with response time measured after an automated voice-activity detector marks the end of speech. This makes comparisons more reproducible, but does not establish identical error rates on noisy phone calls, specialist terms or customer-specific accents.

Microsoft's model overview also lists important omissions: speaker diarisation, word-level timestamps and contextual keyword biasing are not supported. Contact centres may therefore need additional components or a later version.

Two output models, different latency definitions

MAI-Voice-2.1 costs $22 per million characters and lists roughly 550 milliseconds of model inference. Voice-2.1-Flash costs $15 and lists about 45 milliseconds of inference; Microsoft separately claims 150 milliseconds end-to-end for generating 45 seconds of audio. These figures cover different stages and cannot be compared directly with time to the first transcript.

Both voice models support 23 languages, according to the launch post. Microsoft's locale documentation is inconsistent: the announcement says 26 locales, while the product page visibly lists 28 regional entries. The model card describes voice cloning from five to 60 seconds of reference audio; access is gated and tied to consent safeguards.

Pandorex Analysis: a faster audio edge, not the entire agent

Microsoft reduces latency at the input and output edges and lets an application start compute from provisional text. Total conversational delay still depends on end-of-speech detection, the reasoning model, tools, networking and playback. Partial transcripts can also change, so agents should wait for stable segments or confirmation before irreversible actions.

The supported advance is an integrated audio stack with independently validated ASR accuracy. The stronger claim that it automatically produces a delay-free voice agent is not established.

Sources and references

Sources used for the facts and context in this article.

  1. Microsoft AI, 01.10.2026: Our first streaming transcription model debuts at no. 1 on Artificial Analysismicrosoft.ai
  2. Microsoft AI: MAI-Transcribe-2 Modelübersichtmicrosoft.ai
  3. Microsoft AI: MAI-Voice-2.1 Modelübersichtmicrosoft.ai
  4. Microsoft AI: MAI-Voice-2.1 Model Cardmicrosoft.ai
  5. Microsoft Learn: MAI-Voice in Azure Speechlearn.microsoft.com
  6. Artificial Analysis: Speech to Text Streaming Leaderboard und Methodikartificialanalysis.ai

How Pandorex researches and corrects articles

Comments

Sign in to write a comment.

Swipe up
Next Article

Gemini 4 Argon: One Million Output Tokens, but No Public Launch Yet

AI & Chips