Skip to main content

Meta Releases Muse Voice Transcribe Real-Time Model

Meta Superintelligence Labs launches Muse Voice Transcribe, a real-time audio perception model supporting 70+ languages at $3.00 per 1,000 minutes.

AI-written
Inewgen
02 Sep 2026Source: MarkTechPost3 min read (0 views)
Share
Meta Releases Muse Voice Transcribe Real-Time Model

Stock photo for illustration only, not from the actual event

Font size
  • Combines ASR, speaker diarization, and endpointing into a single model
  • Available via Meta Model API at $3.00 per 1,000 audio minutes
  • Supports audio exceeding one hour and 20+ speakers natively

Meta Superintelligence Labs has announced Muse Voice Transcribe, its first real-time audio perception model from the Muse Spark family, designed to collapse three distinct tasks into a single autoregressive model. The system performs streaming ASR, speaker diarization for over 20 speakers, and endpointing in one pass without requiring any post-processing.

As for deployment, the model is currently accessible solely as a hosted API on the Meta Model API under the identifier muse-voice-transcribe-1.0, priced at $3.00 per 1,000 audio minutes ($0.18 per hour). It already powers dictation features in Meta AI for Mac and Muse Code, while no model weights have been publicly released for self-hosting.

3.1%Final-transcript WER
$3.00Price per 1,000 minutes
20+Max speaker diarization

The architecture processes audio arriving in 80ms chunks at 12.5 Hz, transforming each chunk into a single soft token. Listening and writing share a unified decoder loop, eliminating separate alignment stages prone to drift. Meta trains the model using reinforcement learning, combining word error rate and delay rewards multiplicatively to dynamically adjust delay per word according to difficulty.

chromebook notebook computer office desk workspace

Stock photo for illustration only, not from the actual event

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

According to benchmarks on Artificial Analysis as of September 1, 2026, Muse Voice Transcribe records a 3.1% final-transcript WER at 0.16s after the end of speech, outperforming competing systems like Cartesia Ink-2 and ElevenLabs Scribe v2 Realtime. For speaker diarization, it achieves a 17.5% average error rate across AMI-IHM, AMI-SDM, and VoxConverse datasets.

By unifying transcription, diarization, and endpointing into a single autoregressive loop and utilizing reinforcement learning for delay optimization, Meta sets a new Pareto frontier for speed versus accuracy in streaming speech. However, restricting deployment exclusively to a hosted API means organizations with strict on-premise requirements cannot self-host the weights.

The model was trained on more than 70 languages, with 25 extensively verified at launch, and features native code-switching capabilities within and between sentences. Furthermore, it natively handles audio inputs exceeding one hour in length alongside multi-speaker environments without additional post-processing steps.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article