Muse Voice Transcribe Benchmarks
A dated, source-led view of Muse Voice Transcribe performance—without mixing streaming and batch tests or turning launch claims into permanent facts.
Last verified · September 2, 2026
Meta reported that Muse Voice Transcribe ranked first on Artificial Analysis streaming speech-to-text and on public diarization benchmarks. The announcement explicitly dates model inclusion and rankings to September 1, 2026.
- Treat the rank as a launch-day snapshot.
- Link the current leaderboard next to any numerical comparison.
- Keep streaming and non-streaming results in separate tables.
Streaming transcription is a two-axis problem: a system can wait longer to improve its final words, or respond sooner with less context. Meta's adaptive-delay training is designed to move along that trade-off word by word.
Word error rate measures transcription, while diarization error rate measures who spoke when. A trustworthy comparison identifies the datasets, speaker conditions, streaming mode, and whether post-processing was allowed.
We do not present fixed 3.1% WER, 180ms latency, or $0.18/hour as official facts until each value can be verified against a current primary source with matching test conditions.