Meta unveils AI model for real-time multilingual transcription
Meta has unveiled a new AI model capable of transcribing long, real-time conversations involving multiple speakers and languages while identifying who said what.
Called MUSE Voice Transcribe, the model was introduced by Meta Superintelligence Labs on September 1. Meta said it is the company’s first real-time audio perception model, combining speech-to-text transcription, speaker identification and detection of when one speaker stops and another begins.
The model is designed to process conversations involving more than 20 speakers, making it suitable for meetings, interviews, discussions and other extended conversations.
MUSE Voice Transcribe uses streaming automatic speech recognition (ASR), allowing speech to be converted into text as a conversation takes place rather than after the entire recording has been processed. Meta said the system can prioritise speed when speech is clear while spending more time analysing audio when words or context are ambiguous.
The model also supports speaker diarisation, enabling it to separate transcripts by speaker. Its endpointing capability determines when one speaker has finished and another has started, helping maintain the sequence of a conversation.
Meta said the model was trained using data from more than 70 languages and tested on more than 25 languages. It can also detect code-switching, allowing it to handle conversations in which speakers shift between languages.
MUSE Voice Transcribe is currently available through Meta’s Model API and can also be used with Meta AI’s dictation feature for Mac and MUSE Code. Meta has set the API price at $3 per 1,000 minutes of audio.
The technology could be used for live transcription of online meetings and interviews, as well as in education, customer service and voice-based AI applications, particularly where conversations involve multiple speakers or languages.
Leave A Comment