- GPT0%
- MUSE0%
- MINI0%
- SPK0%
- OPUS0%
Perceive Beating AI News Flash: Shortly after the release of Gemini 3.8 Flash, Google achieved a score of 73.7% on DeepSWE v1.1. The DeepSWE official leaderboard currently displays it as 74%±1%, ranking above Claude Opus 5 and GPT-5.6 Sol. Several hours later, Meta announced Muse Spark 1.3, reaching 75.4% on the same DeepSWE suite, pushing the public score up by another 1.7 percentage points.
This test involved the Agent autonomously handling long-term software engineering tasks in 113 real code repositories. Both Google and Meta used the mini-swe-agent, with Google running on high inference mode and Meta on max. Meta's previous generation, Spark 1.2, scored only 55.0%, seeing a direct increase of 20.4 percentage points this time. However, Meta's 75.4% has not yet been included in the DeepSWE official leaderboard.
In the Artificial Analysis Coding Agent Index, the preview-limited Spark 1.3 max received a score of 68, ranking just below Claude Opus 5; the currently publicly available xhigh version scored 64.
Disclaimer: The current content is sourced from third-party perspectives or directly translated by AI from third-party perspectives. CoinEx does not guarantee the authenticity, accuracy, and originality of the content, and it does not constitute any investment advice from CoinEx. The prices of cryptocurrencies are highly volatile, please be aware of the potential risks.
- CoinsPrice24H Change