- GPT0%
- ASR0%
- MUSE0%
Perceive Beating AI News Flash: Meta Superintelligence Labs has released Muse Voice Transcribe. It performs real-time speech-to-text conversion, distinguishes between different speakers, and determines when a user has finished speaking, all in a single model. It can process audio files of up to 1 hour in length and differentiate between over 20 individuals.
In an English real-time transcription test by Artificial Analysis, its WER (Word Error Rate, lower is better) was 3.1%. GPT Live Transcribe scored 3.9%, and Gemini 3.5 Transcribe Live scored 4.0%. Muse Voice Transcribe can provide the final text approximately 0.16 seconds after the speaker finishes.
It does not assign the same waiting time to all words. The model reads 80ms of audio each time and decides whether to continue listening or start outputting text. Simple words can be transcribed more quickly, while difficult words are listened to for more context before confirmation. Meta uses reinforcement learning to optimize both accuracy and latency simultaneously.
The model has been trained on over 70 languages, with 25 languages undergoing intensive validation. It can also directly recognize mixed Chinese and English speech in a single sentence. Currently integrated into Meta AI for Mac and Muse Code, it has also opened up the Meta Model API. The pricing is $3 per 1000 minutes, which is equivalent to $0.18 per hour.
免責聲明:當前內容均來自第三方觀點或由AI直接翻譯第三方觀點,CoinEx不保證內容的真實性、準確性和原創性,不構成CoinEx相關的任何投資建議。數字資產價格波動劇烈,請注意潛在風險。
- 幣種價格24H漲跌