Benchmark methodology

Muse Voice Transcribe Benchmarks

A dated, source-led view of Muse Voice Transcribe performance—without mixing streaming and batch tests or turning launch claims into permanent facts.

Last verified · September 2, 2026

01
What Meta reported at launch

Meta reported that Muse Voice Transcribe ranked first on Artificial Analysis streaming speech-to-text and on public diarization benchmarks. The announcement explicitly dates model inclusion and rankings to September 1, 2026.

  • Treat the rank as a launch-day snapshot.
  • Link the current leaderboard next to any numerical comparison.
  • Keep streaming and non-streaming results in separate tables.
02
Accuracy and delay must be read together

Streaming transcription is a two-axis problem: a system can wait longer to improve its final words, or respond sooner with less context. Meta's adaptive-delay training is designed to move along that trade-off word by word.

03
Diarization is a separate measurement

Word error rate measures transcription, while diarization error rate measures who spoke when. A trustworthy comparison identifies the datasets, speaker conditions, streaming mode, and whether post-processing was allowed.

04
Numbers we do not publish yet

We do not present fixed 3.1% WER, 180ms latency, or $0.18/hour as official facts until each value can be verified against a current primary source with matching test conditions.