- GPT0%
- ASR0%
- MUSE0%
Perceive Beating AI News Flash: Meta Superintelligence Labs has released Muse Voice Transcribe. It performs real-time speech-to-text conversion, distinguishes between different speakers, and determines when a user has finished speaking, all in a single model. It can process audio files of up to 1 hour in length and differentiate between over 20 individuals.
In an English real-time transcription test by Artificial Analysis, its WER (Word Error Rate, lower is better) was 3.1%. GPT Live Transcribe scored 3.9%, and Gemini 3.5 Transcribe Live scored 4.0%. Muse Voice Transcribe can provide the final text approximately 0.16 seconds after the speaker finishes.
It does not assign the same waiting time to all words. The model reads 80ms of audio each time and decides whether to continue listening or start outputting text. Simple words can be transcribed more quickly, while difficult words are listened to for more context before confirmation. Meta uses reinforcement learning to optimize both accuracy and latency simultaneously.
The model has been trained on over 70 languages, with 25 languages undergoing intensive validation. It can also directly recognize mixed Chinese and English speech in a single sentence. Currently integrated into Meta AI for Mac and Muse Code, it has also opened up the Meta Model API. The pricing is $3 per 1000 minutes, which is equivalent to $0.18 per hour.
면책 조항: 현재 콘텐츠는 제3자 관점에서 제공되거나 제3자 관점에서 AI가 직접 번역한 것입니다. CoinEx는 콘텐츠의 진위성, 정확성, 독창성을 보장하지 않으며 CoinEx의 투자 조언으로 간주하지 않습니다. 암호화폐 가격은 변동성이 크므로 잠재적인 위험에 유의하시기 바랍니다.
- 코인가격24시간 변동