Meta prices Muse Voice Transcribe at $0.18 an hour, with real-time diarization for 20+ speakers: a steal for enterprises?
Meta launched Muse Voice Transcribe, offering real-time voice transcription, endpoint detection, and speech segmentation for over 20 speakers at a price of $0.18 per hour. Developed by Meta’s Super Intelligent Lab, this model supports long audio recordings (over one hour), seamless multilingual code switching, language and keyword bias, and speech segmentation without the need for separate post-processing pipelines. The model was trained in 70 languages, with 25 of them having extensive verification. Despite the high number of speakers, Meta’s core proposition is to integrate high-capacity real-time segmentation with low-latency transcription, endpoint detection, multilingual code switching, and aggressive API pricing into a single model, which may be even more important for enterprise developers building conference systems, call analytics, real-time assistants, or environmental AI. Segmentation is gradually becoming a core component of the voice stack, and Muse directly embeds speaker attribution into the autoregressive multimodal architecture, applying adaptive latency processing when audio reaches 80 milliseconds in length…