US-based technology giant Meta has introduced Muse Voice Transcribe, its first real-time audio-perception model, with native support for five major Indian languages and advanced streaming capabilities.
Meta introduces Muse Voice Transcribe for real-time speech
Muse Voice Transcribe — developed by Meta Superintelligence Labs — delivers streaming transcription, speaker separation across over 20 voices in hour‑plus recordings, native code‑switching and diarisation from a single model with no post‑processing step, the company said in a statement.
Muse Voice Transcribe is trained across over 70 languages spoken in multiple countries with 25 validated at launch.
Muse Voice Transcribe targets faster and more accurate transcription
“It ranks first on the Artificial Analysis streaming speech-to-text leaderboard as of September 1, 2026,” the statement added.
Available via the Meta Model API and already running dictation in Meta AI for Mac and Muse Code, Muse Voice Transcribe delivers real‑time automatic speech recognition, diarisation with over 20 speakers and endpointing.
It is multilingual with seamless code-switching and improves accuracy with language, keyword, and context biasing.
Meta uses adaptive delay to balance speed and accuracy
“The longer the model waits to predict, the more accurate the transcript, but the higher the latency. Muse Voice Transcribe has “adaptive delay,” dynamically changing delay for each word based on difficulty,” the statement noted.
This is enabled with reinforcement learning (RL), where word error rate (WER) reward and a delay reward are combined multiplicatively.
How Meta’s Muse model processes audio
“Muse Voice Transcribe is an autoregressive multimodal model from the Muse Spark family,” Meta said.
It explained that the audio is processed in 80 ms chunks (12.5 Hz), each of which is transformed into a single soft token. At each audio chunk, the model decides to either continue listening to the next audio chunk or emit a text token.
With adaptive delay, Muse Voice Transcribe achieves the Pareto front on speed-accuracy trade-off measured by time to final transcription, the company said.
Read More:
Google Restricts Meta Gemini AI Access as Computing Capacity Crunch Delays AI Projects
Meta’s AI model hacked external system during cybersecurity test










