AuraTracer智迹闻
中文

EVENT DOSSIER

Meta prices Muse Voice Transcribe at $0.18 an hour, with real-time diarization for 20+ speakers: a steal for enterprises?

2026-09-03 05:32 Models 🔥 38.3 heat score
1sources
1days unfolding
38.3heat score
5mentions
SummaryAI generated

On September 2, 2026, Meta launched Muse Voice Transcribe, developed by Super Intelligent Lab. This model offers real-time voice transcription, endpoint detection, and word segmentation for over 20 speakers at a price of $0.18 per hour. It supports processing of long audio files lasting more than an hour, seamless switching between multiple languages, and is trained in 70 languages (of which 25 have been widely verified). Muse integrates speaker attribution directly into its autoregressive multimodal architecture, enabling adaptive delay processing when audio reaches 80 milliseconds in length, allowing for high-capacity real-time transcription without the need for separate post-processing pipelines.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
Amazon TranscribeMetaMuse Voice TranscribeSpeechmaticsSuperintelligence Labs

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Amazon Transcribe × Meta1Amazon Transcribe × Mus…1Amazon Transcribe × Spe…1Amazon Transcribe × Sup…1Meta × Muse Voice Trans…1Meta × Speechmatics1

SignalsSIGNALS

Keyword heat
  • Meta1
  • Muse Voice Transcribe1
  • Superintelligence Labs1
  • Speechmatics1
  • Amazon Transcribe1

All reports (1)SOURCES

V VentureBeat AI en 2026-09-03 05:32

Meta prices Muse Voice Transcribe at $0.18 an hour, with real-time diarization for 20+ speakers: a steal for enterprises?

Meta launched Muse Voice Transcribe, offering real-time voice transcription, endpoint detection, and speech segmentation for over 20 speakers at a price of $0.18 per hour. Developed by Meta’s Super Intelligent Lab, this model supports long audio recordings (over one hour), seamless multilingual code switching, language and keyword bias, and speech segmentation without the need for separate post-processing pipelines. The model was trained in 70 languages, with 25 of them having extensive verification. Despite the high number of speakers, Meta’s core proposition is to integrate high-capacity real-time segmentation with low-latency transcription, endpoint detection, multilingual code switching, and aggressive API pricing into a single model, which may be even more important for enterprise developers building conference systems, call analytics, real-time assistants, or environmental AI. Segmentation is gradually becoming a core component of the voice stack, and Muse directly embeds speaker attribution into the autoregressive multimodal architecture, applying adaptive latency processing when audio reaches 80 milliseconds in length…