- OPUS0%
Dynamic Beating AI News: AI Agent company Einsia has released the SWE Refactor Bench, specifically designed to test if a Coding Agent can autonomously refactor an entire codebase. The benchmark consists of 20 tasks from real projects such as SQLite, zlib, libsodium, GraphHopper, involving tasks like translating C to Rust, migrating from Maven to Gradle, and porting SQLite to WASI. Each task was allocated 6 to 30 hours for completion.
The first evaluation stage checks if the old tech stack has indeed been replaced to prevent the Agent from carrying over legacy code. The second stage runs over 130,000 fixed checks. Finally, 6 independent Coding Agents are deployed, each spending 1 hour to specifically search for hidden bugs.
A total of 520 runs were conducted using 8 cutting-edge models and 26 configurations. Out of these, 340 runs successfully completed the migration, with 88 runs passing all pre-defined tests. However, in 60 cases, hidden bugs were discovered in the final round. Only 28 runs successfully passed all stages, resulting in a pass rate of 5.4%. Out of the 20 tasks, 13 were left incomplete. The xhigh configuration of Claude Opus 5 performed the best, completing all 20 tasks.
The current Coding Agents are now capable of large-scale code modifications. However, achieving a state where "the entire system is modified without any errors" is still a distant goal.
Disclaimer: The current content is sourced from third-party perspectives or directly translated by AI from third-party perspectives. CoinEx does not guarantee the authenticity, accuracy, and originality of the content, and it does not constitute any investment advice from CoinEx. The prices of cryptocurrencies are highly volatile, please be aware of the potential risks.
- CoinsPrice24H Change