Meta AI Research · Released Sep 01, 2026

Muse Voice Transcribe real-time speech-to-text guide

Find Meta's muse-voice-transcribe-1.0 model, API pricing, language coverage, live transcription, speaker diarization, and endpointing facts. This is an independent guide and is not affiliated with Meta.

Unofficial disclosure: musevoicetranscribe.site is not an official Meta, Meta AI, or Meta Superintelligence Labs website.

80 ms
audio chunks
20+
speakers
70+
training languages
> 1 hr
native long context
Direct answer

What is Muse Voice Transcribe?

Muse Voice Transcribe is a real-time audio perception model from Meta Superintelligence Labs for streaming speech recognition, speaker diarization, and endpointing. It is designed to work while audio is still arriving, not only after a full recording is uploaded.

Public Meta materials describe access through Meta Model API, Meta AI for Mac, and Muse Code. This site summarizes those public facts and provides an independent browser transcription experience.

Live transcription

Turn your voice into text now

Speak naturally, see words appear in real time, then copy or download a clean transcript.

Live browser transcription
Speak naturally, watch words appear in real time, then copy or download your transcript.
Ready
Elapsed time: 00:00
Your live transcript will appear here. You can also load a prepared example to preview the transcript controls.

Select a language, then allow microphone access to begin.

Officially documented

One model, three live signals

Meta combines transcription, speaker identity, and turn boundaries in one streaming model rather than stitching together separate post-processing systems.

Streaming ASR
Processes audio as it arrives instead of waiting for a complete file.
20+ speakers
Native speaker diarization for long, multi-person conversations.
Endpointing
Detects speech onset and when a conversational turn is complete.
Multilingual
Trained on 70+ languages, with 25 extensively verified at launch.
Model design

Listen longer, or emit now?

For every 80ms audio chunk, the model chooses whether to keep listening for context or emit text. Adaptive delay changes that wait dynamically based on difficulty.

80ms audio chunk

A compact audio representation enters the stream.

Listen or emit

The autoregressive model asks for more audio or returns text.

Adaptive delay

Harder words get more context; easier words arrive sooner.

Evidence, dated

Rankings move. Sources should not.

Meta reported Muse Voice Transcribe ranked first on Artificial Analysis streaming speech-to-text and public diarization benchmarks at launch. That claim is explicitly dated September 1, 2026.

See benchmark methodology
FAQ

Quick checks on access, pricing, and limits

Is this an official Meta website?

No. This is an independent guide and browser transcription utility. Treat Meta AI Research and current Meta Model API documentation as the source of truth for model capabilities.

Is Muse Voice Transcribe free?

This site does not sell API credits. Meta lists Muse Voice Transcribe on Meta Model API at $0.18 per audio hour; Meta AI for Mac and Muse Code access should not be read as API billing documentation.

How do I use Muse Voice Transcribe?

Developers should evaluate muse-voice-transcribe-1.0 through Meta Model API docs. Non-developers can follow the Meta AI for Mac dictation path where Meta says the model is already used.

Does it work on Mac?

Meta says Muse Voice Transcribe powers dictation in Meta AI for Mac and Muse Code. Region, account, and product availability should be checked in Meta's current product pages.

Which languages are supported?

Meta says the model is trained on 70+ languages, with 25 extensively verified at launch, and supports multilingual speech and in-sentence code-switching.

How is it different from Whisper?

Muse Voice Transcribe is positioned for real-time streaming ASR with diarization and endpointing in one model. Whisper is often used for batch transcription and self-hosted workflows; the right choice depends on latency, privacy, cost, and deployment needs.

Research sources

Start with primary sources, not marketing numbers.

Model capabilities are grounded in Meta AI Research and current developer documentation; third-party measurements are labeled and dated.

Open the official announcement
Muse Voice Transcribe: Live Demo, Benchmarks & Guide