- GPT0%
- ASR0%
- MUSE0%
Perceive Beating AI News Flash: Meta Superintelligence Labs has released Muse Voice Transcribe. It performs real-time speech-to-text conversion, distinguishes between different speakers, and determines when a user has finished speaking, all in a single model. It can process audio files of up to 1 hour in length and differentiate between over 20 individuals.
In an English real-time transcription test by Artificial Analysis, its WER (Word Error Rate, lower is better) was 3.1%. GPT Live Transcribe scored 3.9%, and Gemini 3.5 Transcribe Live scored 4.0%. Muse Voice Transcribe can provide the final text approximately 0.16 seconds after the speaker finishes.
It does not assign the same waiting time to all words. The model reads 80ms of audio each time and decides whether to continue listening or start outputting text. Simple words can be transcribed more quickly, while difficult words are listened to for more context before confirmation. Meta uses reinforcement learning to optimize both accuracy and latency simultaneously.
The model has been trained on over 70 languages, with 25 languages undergoing intensive validation. It can also directly recognize mixed Chinese and English speech in a single sentence. Currently integrated into Meta AI for Mac and Muse Code, it has also opened up the Meta Model API. The pricing is $3 per 1000 minutes, which is equivalent to $0.18 per hour.
Disclaimer: The current content is sourced from third-party perspectives or directly translated by AI from third-party perspectives. CoinEx does not guarantee the authenticity, accuracy, and originality of the content, and it does not constitute any investment advice from CoinEx. The prices of cryptocurrencies are highly volatile, please be aware of the potential risks.
- CoinsPrice24H Change