On September 6, 2026, Meta released a real-time audio transcription model called Muse Voice Transcribe. This model can process voice segments in 80 milliseconds, distinguish between different speakers, and detect sentence boundaries. According to Artificial Analysis, this model offers the lowest price and most accurate streaming transcription service in the market. Meta plans to use this model as a fundamental component for building personal AI agents capable of continuously listening to real conversations, and it is expected to be applied in devices such as smart glasses.
Meta released Muse Voice Transcribe, a real-time transcription model that processes voice segments in 80 milliseconds, distinguishes speakers, and detects sentence boundaries. According to Artificial Analysis, this model offers the most accurate streaming transcription service at the lowest price in the market. Meta views it as the fundamental building block for personal AI agents that can listen to real conversations through devices such as smart glasses.