- GOOGLX0%
- GPT0%
Dynamic Insight AI News Flash: Google has released Gemini 3.5 Transcribe, which is currently the most accurate Google speech-to-text model. This is also the first time Google has included a Smart mode in a dedicated transcription model. The model is available in both real-time and recording transcription versions, and the Gemini API is now in open beta. In Artificial Analysis evaluations, the non-streaming WER (Word Error Rate, lower is better) is 2.6% for the recording version and 4.0% for the real-time version. The final transcription latency has been reduced by 70% compared to the previous generation Chirp 3.
The Smart mode is no longer just about converting speech into text word by word. For example, if you say, "Let's meet on Tuesday, no, Wednesday," it will directly keep "Wednesday"; filler words like "Um, Ah" are automatically removed. Additionally, it can organize colloquial speech into paragraphs, lists, dates, and numbers. When verbatim meeting records are needed, the default word-for-word transcription mode can still be used.
The model supports over 85 languages and on-the-fly language switching, and allows the inclusion of custom professional vocabulary. The recording transcription is priced at around $0.005 per minute, which is $5 per 1000 minutes; the real-time version costs approximately $0.009 per minute. On a standalone leaderboard, it is more accurate than OpenAI's GPT Transcribe at 3.3%, but has not yet surpassed ElevenLabs Scribe v2, which has a WER of 2.2% at around $3.67 per 1000 minutes.
This model has already been used in Android's Rambler and macOS's Gemini, and will soon be integrated into Chrome. This will allow users to speak and type directly into the webpage input box.
免責事項:現在のコンテンツは第三者の視点に基づくもの、または第三者の視点からAIが直接翻訳したものです。CoinExはコンテンツの信頼性、正確性、独創性を保証するものではなく、CoinExからの投資アドバイスを構成するものではありません。暗号資産の価格変動は急激に変動します。潜在的なリスクにご注意ください。
- コインリスト価格24時間価格変動