- GPT0%
- MUSE0%
- MINI0%
- SPK0%
- OPUS0%
Perceive Beating AI News Flash: Shortly after the release of Gemini 3.8 Flash, Google achieved a score of 73.7% on DeepSWE v1.1. The DeepSWE official leaderboard currently displays it as 74%±1%, ranking above Claude Opus 5 and GPT-5.6 Sol. Several hours later, Meta announced Muse Spark 1.3, reaching 75.4% on the same DeepSWE suite, pushing the public score up by another 1.7 percentage points.
This test involved the Agent autonomously handling long-term software engineering tasks in 113 real code repositories. Both Google and Meta used the mini-swe-agent, with Google running on high inference mode and Meta on max. Meta's previous generation, Spark 1.2, scored only 55.0%, seeing a direct increase of 20.4 percentage points this time. However, Meta's 75.4% has not yet been included in the DeepSWE official leaderboard.
In the Artificial Analysis Coding Agent Index, the preview-limited Spark 1.3 max received a score of 68, ranking just below Claude Opus 5; the currently publicly available xhigh version scored 64.
Disclaimer: Konten ini berasal dari pihak lain atau diterjemahkan oleh AI dari pihak lain. CoinEx tidak menjamin konten ini benar, asli, atau akurat, dan tidak memberikan saran investasi. Harga aset kripto sangat tidak stabil, jadi harap berhati-hati terhadap risiko yang ada.
- KriptoHargaPerubahan 24J