- MODE0%
- MYTH0%
- OPUS0%
According to Watchful AI, Anthropic disclosed in the Claude Opus 5 system card that the model is almost immune to prompt injection attacks in the browser agent scenario, with none of the 129 test scenarios being compromised. This breakthrough is significant as OpenAI publicly acknowledged last December that prompt injection may never be fully resolved. In Gray Swan's general prompt injection benchmark test, after 15 attacks, Opus 5 achieved a success rate of only 2.0%, significantly better than the previous generation Opus 4.8 at 5.5%, as well as outperforming Mythos 5 at 2.6% and Fable 5 at 2.8%.
Prompt injection is considered the most serious security vulnerability faced by AI agents—attackers manipulate input content such as hidden text embedded in a webpage to bypass model instructions and perform malicious operations. A zero-attack rate is only achieved in products like Claude Cowork with Auto Mode enabled. This achievement marks a transition in AI agent security from "ongoing patches" to "structural defense."
Click on the original article link below to join the Watchful AI · Feishu AI News Channel and receive real-time monitoring of global AI trends and news 24/7.
免責事項:現在のコンテンツは第三者の視点に基づくもの、または第三者の視点からAIが直接翻訳したものです。CoinExはコンテンツの信頼性、正確性、独創性を保証するものではなく、CoinExからの投資アドバイスを構成するものではありません。暗号資産の価格変動は急激に変動します。潜在的なリスクにご注意ください。
- コインリスト価格24時間価格変動