Models#Speech to text

Microsoft ships its first streaming transcription model — and it debuts at No.1 on both boards

MAI-Transcribe-2-Streaming takes No.1 on Artificial Analysis for final and partial transcripts: 2.50% WER, 0.13s latency, 60 languages, $0.54/hour promotional.

Microsoft art: quotation marks flying over a blue field with Fast, accurate, low-cost text

Microsoft AI shipped MAI-Transcribe-2-Streaming on October 1 — the company’s first streaming speech-to-text model, debuting at No.1 on Artificial Analysis’s streaming leaderboard for both final-transcript accuracy and partial-transcript latency. Two voice models, MAI-Voice-2.1 and Voice-2.1-Flash, shipped alongside.

The facts

  • Board results: first place among 28-38 streaming STT models; 2.50% word error rate (vs Grok Voice Transcribe 2.0 Streaming at 2.73% and ElevenLabs Scribe v2 Realtime at 3.59%).
  • Latency: full transcription 0.13s after speech ends; first partial text ~100ms after speech starts, text at ~320ms.
  • Languages: 60, with continuous automatic language detection.
  • Availability: Microsoft Foundry, MAI Playground and OpenRouter; promotional pricing of $0.54/hour; an Azure Speech streaming tier is not mentioned.
  • Sourcing note: Microsoft’s post and IT之家’s report agree; the leaderboard numbers come from Artificial Analysis’s September 28 evaluation.

Why the realtime voice layer needed this

Until now the streaming-accuracy-versus-latency trade had no citable winner — every vendor benchmarked its own scenario. A third-party double-first gives voice-agent teams (meetings, support, live translation) a fixed reference: 2.50% at 0.13s, with the caveat that this measures end-of-speech to full text; live captions care about partial latency, which is a different line on the same board.

Where it sits in the MAI line

This extends Microsoft’s in-house MAI line from text and image into voice, complementing the Copilot product-line restructure from the same week — a “work entry point” needs its own ears and voice. The replace-Whisper route is getting more real with each release.

Editorial take

At $0.54 per hour promotional, streaming transcription crosses into “cost is not the question” territory — an hour-long meeting transcribes for under four RMB. Voice-product teams should run this through their own eval sets this week: measure what 2.50% means on real noisy audio before believing a leaderboard number.

The leaderboard context

Artificial Analysis’s streaming board is the first to normalize the awkward trade in live transcription: partial latency (how fast first words appear), final latency (how fast the settled text lands) and accuracy, across providers. Past leaders held one axis; MAI-Transcribe-2-Streaming’s double first at a promotional price is what changes procurement math. The caveats: leaderboard audio is clean-ish, so teams with domain jargon or noisy channels should still validate, and the Azure Speech streaming tier has not appeared — the model lives on Microsoft Foundry, MAI Playground and OpenRouter.

The competitive picture

ElevenLabs (Scribe v2 Realtime), Grok Voice and Deepgram now trail on the public board, and OpenAI’s Whisper heritage is notably absent from the streaming top ranks — Microsoft’s MAI line is attacking exactly the gap OpenAI left while those teams focus elsewhere. For OpenRouter availability matters too: a Microsoft model on a third-party router is a new distribution posture for the MAI line.

What it means for voice-agent builders

The end-of-speech 0.13s figure suits turn-taking agents (voice assistants that must know when the user stopped talking); partial latency matters for captions and translation overlays. At $0.54/hour promotional, the cost floor for “transcribe everything” products — always-on meeting capture, call analytics — drops below the storage cost of the audio itself, which is the real signal: transcription is becoming a free primitive, and the value moves to what you do with the transcript.

One more integration note: with continuous language detection across 60 languages, mixed-language meetings no longer need pre-routing — the model switches mid-stream. That single feature retires a whole category of language-detection preprocessing in call-analytics pipelines, which is often where those projects quietly died.

For teams that already run MAI models, the streaming release also completes the realtime voice loop Microsoft needs for Copilot Voice — one vendor supplying hearing, speech and the agent brain, with the integration tax paid once. That bundling is the quiet strategy behind three model releases in one day.