📣 Send us your press release
Site updates every 15 minutes
Technology

Microsoft launches real-time speech-to-text model

Microsoft has released MAI-Transcribe-2-Streaming, its first AI model for continuous, real-time speech-to-text transcription in 60 languages. The model offers fast processing and high accuracy.

1 October 2026
Microsoft launches real-time speech-to-text model
Image is an AI-generated illustration

Microsoft announced the launch of MAI-Transcribe-2-Streaming, its first AI model for continuous, real-time speech-to-text transcription. The model generates text as speech occurs and supports 60 languages with automatic language detection.

The new model is designed to enhance real-time applications, such as customer service bots and live captioning systems. It can understand user intent and initiate actions before a speaker finishes their sentence, providing a more immediate experience. The model delivers initial text segments rapidly and refines them for stability.

According to Microsoft's internal evaluations, the model can display the first text results approximately 320 milliseconds after speech begins. In independent tests by Artificial Analysis, the model achieved a final transcription latency of 0.13 seconds and a word error rate of 2.50%, ranking it first among 28 tested streaming speech-to-text models.

MAI-Transcribe-2-Streaming is available to developers through channels like Microsoft Foundry and MAI Playground. It is currently offered at an introductory price of $0.54 per hour. Microsoft previously released the non-streaming MAI-Transcribe-2 model.

Original source: ithome.com