- OPUS0%
Dynamic Beating AI News: AI Agent company Einsia has released the SWE Refactor Bench, specifically designed to test if a Coding Agent can autonomously refactor an entire codebase. The benchmark consists of 20 tasks from real projects such as SQLite, zlib, libsodium, GraphHopper, involving tasks like translating C to Rust, migrating from Maven to Gradle, and porting SQLite to WASI. Each task was allocated 6 to 30 hours for completion.
The first evaluation stage checks if the old tech stack has indeed been replaced to prevent the Agent from carrying over legacy code. The second stage runs over 130,000 fixed checks. Finally, 6 independent Coding Agents are deployed, each spending 1 hour to specifically search for hidden bugs.
A total of 520 runs were conducted using 8 cutting-edge models and 26 configurations. Out of these, 340 runs successfully completed the migration, with 88 runs passing all pre-defined tests. However, in 60 cases, hidden bugs were discovered in the final round. Only 28 runs successfully passed all stages, resulting in a pass rate of 5.4%. Out of the 20 tasks, 13 were left incomplete. The xhigh configuration of Claude Opus 5 performed the best, completing all 20 tasks.
The current Coding Agents are now capable of large-scale code modifications. However, achieving a state where "the entire system is modified without any errors" is still a distant goal.
免責聲明:當前內容均來自第三方觀點或由AI直接翻譯第三方觀點,CoinEx不保證內容的真實性、準確性和原創性,不構成CoinEx相關的任何投資建議。數字資產價格波動劇烈,請注意潛在風險。
- 幣種價格24H漲跌