- GPT0%
- MUSE0%
- MINI0%
- SPK0%
- OPUS0%
Perceive Beating AI News Flash: Shortly after the release of Gemini 3.8 Flash, Google achieved a score of 73.7% on DeepSWE v1.1. The DeepSWE official leaderboard currently displays it as 74%±1%, ranking above Claude Opus 5 and GPT-5.6 Sol. Several hours later, Meta announced Muse Spark 1.3, reaching 75.4% on the same DeepSWE suite, pushing the public score up by another 1.7 percentage points.
This test involved the Agent autonomously handling long-term software engineering tasks in 113 real code repositories. Both Google and Meta used the mini-swe-agent, with Google running on high inference mode and Meta on max. Meta's previous generation, Spark 1.2, scored only 55.0%, seeing a direct increase of 20.4 percentage points this time. However, Meta's 75.4% has not yet been included in the DeepSWE official leaderboard.
In the Artificial Analysis Coding Agent Index, the preview-limited Spark 1.3 max received a score of 68, ranking just below Claude Opus 5; the currently publicly available xhigh version scored 64.
免責事項:現在のコンテンツは第三者の視点に基づくもの、または第三者の視点からAIが直接翻訳したものです。CoinExはコンテンツの信頼性、正確性、独創性を保証するものではなく、CoinExからの投資アドバイスを構成するものではありません。暗号資産の価格変動は急激に変動します。潜在的なリスクにご注意ください。
- コインリスト価格24時間価格変動