Meta launches Muse Voice Transcribe for real-time speech transcription
Meta has entered the speech-to-text market with Muse Voice Transcribe, a model capable of real-time transcription and speaker diarization for over 20 participants. It is priced at $0.18 per hour.

Meta has launched Muse Voice Transcribe, entering the competitive real-time speech-to-text market. The new model combines streaming transcription, endpoint detection, and speaker diarization for more than 20 speakers, priced at $0.18 per hour of processed audio.
Developed by Meta Superintelligence Labs, Muse is designed for real-time processing, rather than waiting for recordings to finish. Meta states the model supports long audio exceeding an hour, seamless multilingual code-switching, language and keyword biasing, and diarization without a separate post-processing pipeline. The model was trained across more than 70 languages, with 25 extensively validated for the initial release.
While the 20-plus speaker capacity is substantial, it is not a market record. Competitors like Speechmatics claim support for up to 100 speakers, and Amazon Transcribe supports up to 30. However, Muse offers a combination of high-capacity real-time diarization, low latency, multilingual capabilities, and aggressive API pricing within a single model.
This combination is significant for enterprise developers building meeting systems, call analytics, or ambient AI applications. Accurate speaker attribution is crucial as transcripts feed downstream AI systems, preventing misattributions that can undermine corporate records or compliance workflows.
At $0.18 per hour, Muse presents a competitive pricing structure. Meta indicates this rate applies to actual processed audio and is the same for both streaming and non-streaming transcription, with zero-data-retention processing offered at parity.