← Latest briefing

Technology

Meta launches Muse Voice Transcribe audio perception model

The new model provides real-time multilingual dictation, speaker separation, and developer API access.

The short version

  • Meta released Muse Voice Transcribe, a real-time speech recognition model capable of handling over 70 languages and speaker diarization.
  • The tool integrates into Mac operating systems to enable system-wide dictation by holding the Fn key.
  • Developers can access the technology via API for $3 per 1,000 audio-minutes.

Key facts

  • Meta launched Muse Voice Transcribe, integrating real-time speech recognition, speaker diarization, and endpointing into a single process.[9to5Mac]
  • The model was trained on more than 70 languages, with 25 validated at launch, and supports native code-switching and recordings over an hour long.[9to5Mac]
  • The system utilizes an adaptive delay mechanism that adjusts context listening times based on speech difficulty to balance speed and accuracy.[9to5Mac]
  • The model is offered via the Meta Model API at a rate of $3 per 1,000 audio-minutes and powers dictation features in Meta AI for Mac and Muse Code.[9to5Mac]

Sources