RADAR ·
Meta releases a real-time transcription model priced at $0.18 per hour
Meta's Superintelligence Labs has released Muse Voice Transcribe, a live transcription model that processes audio in 80-millisecond chunks and decides per word how long to keep listening before committing to text. The same model separates more than 20 speakers and marks sentence boundaries. It supports over 70 languages; 25 were tested in depth.
In an independent evaluation by Artificial Analysis, the model reached a 3.1 percent word error rate on English with a 0.16-second delay after the speaker stops, ahead of ElevenLabs, AssemblyAI and Cartesia on accuracy. The price is $0.18 per hour of audio, below competing services. Meta presents the model as a building block for assistants that listen to conversations through its camera glasses. The weights are not released.
Source: The Decoder