Microsoft AI Releases Faster, Cheaper Speech Recognition Model
Microsoft AI has introduced MAI-Transcribe-2, a speech recognition model that the company claims is faster, more accurate, and cheaper than competing solutions from OpenAI, Google, and ElevenLabs. The new model is priced at 10 cents per hour of audio.

Microsoft AI released its MAI-Transcribe-2 speech recognition model on Thursday, a development the company states offers superior speed, accuracy, and cost-effectiveness compared to offerings from OpenAI, Google, and ElevenLabs. The model is priced at $0.10 per hour of audio processed.
This pricing represents a significant reduction from the previous generation, which cost $0.36 per hour. For businesses processing substantial audio volumes, such as 100,000 hours annually, this price cut could translate to savings of tens of thousands of dollars per year.
The release aligns with Microsoft's broader strategy of developing proprietary, frontier-class AI models across various modalities. Speech recognition is an area where this strategy has seen rapid advancement. MAI-Transcribe-2 expands language support to 60 languages, up from 43 in its predecessor, and is engineered to handle challenging audio conditions common in business environments.
Key features for enterprise users include speaker diarization, word-level timestamps, and keyword biasing for improved accuracy with domain-specific jargon. The model also offers automatic language identification and handles code-switching within conversations. Previously, these capabilities often required separate, premium services.
Microsoft cites benchmark results, including a number one ranking on the FLEURS benchmark across 60 languages and a second-place position on the Artificial Analysis accuracy-latency leaderboard. The company also claims significant speed advantages over competing models, enabling the model's competitive pricing and efficient performance.