- GOOGLX0%
- GPT0%
Dynamic Insight AI News Flash: Google has released Gemini 3.5 Transcribe, which is currently the most accurate Google speech-to-text model. This is also the first time Google has included a Smart mode in a dedicated transcription model. The model is available in both real-time and recording transcription versions, and the Gemini API is now in open beta. In Artificial Analysis evaluations, the non-streaming WER (Word Error Rate, lower is better) is 2.6% for the recording version and 4.0% for the real-time version. The final transcription latency has been reduced by 70% compared to the previous generation Chirp 3.
The Smart mode is no longer just about converting speech into text word by word. For example, if you say, "Let's meet on Tuesday, no, Wednesday," it will directly keep "Wednesday"; filler words like "Um, Ah" are automatically removed. Additionally, it can organize colloquial speech into paragraphs, lists, dates, and numbers. When verbatim meeting records are needed, the default word-for-word transcription mode can still be used.
The model supports over 85 languages and on-the-fly language switching, and allows the inclusion of custom professional vocabulary. The recording transcription is priced at around $0.005 per minute, which is $5 per 1000 minutes; the real-time version costs approximately $0.009 per minute. On a standalone leaderboard, it is more accurate than OpenAI's GPT Transcribe at 3.3%, but has not yet surpassed ElevenLabs Scribe v2, which has a WER of 2.2% at around $3.67 per 1000 minutes.
This model has already been used in Android's Rambler and macOS's Gemini, and will soon be integrated into Chrome. This will allow users to speak and type directly into the webpage input box.
면책 조항: 현재 콘텐츠는 제3자 관점에서 제공되거나 제3자 관점에서 AI가 직접 번역한 것입니다. CoinEx는 콘텐츠의 진위성, 정확성, 독창성을 보장하지 않으며 CoinEx의 투자 조언으로 간주하지 않습니다. 암호화폐 가격은 변동성이 크므로 잠재적인 위험에 유의하시기 바랍니다.
- 코인가격24시간 변동