- OPUS0%
Dynamic Beating AI News: AI Agent company Einsia has released the SWE Refactor Bench, specifically designed to test if a Coding Agent can autonomously refactor an entire codebase. The benchmark consists of 20 tasks from real projects such as SQLite, zlib, libsodium, GraphHopper, involving tasks like translating C to Rust, migrating from Maven to Gradle, and porting SQLite to WASI. Each task was allocated 6 to 30 hours for completion.
The first evaluation stage checks if the old tech stack has indeed been replaced to prevent the Agent from carrying over legacy code. The second stage runs over 130,000 fixed checks. Finally, 6 independent Coding Agents are deployed, each spending 1 hour to specifically search for hidden bugs.
A total of 520 runs were conducted using 8 cutting-edge models and 26 configurations. Out of these, 340 runs successfully completed the migration, with 88 runs passing all pre-defined tests. However, in 60 cases, hidden bugs were discovered in the final round. Only 28 runs successfully passed all stages, resulting in a pass rate of 5.4%. Out of the 20 tasks, 13 were left incomplete. The xhigh configuration of Claude Opus 5 performed the best, completing all 20 tasks.
The current Coding Agents are now capable of large-scale code modifications. However, achieving a state where "the entire system is modified without any errors" is still a distant goal.
Descargo de responsabilidad: El contenido actual proviene de perspectivas de terceros o es traducido directamente por IA a partir de perspectivas de terceros. CoinEx no garantiza la autenticidad, exactitud u originalidad del contenido, y no constituye ningún consejo de inversión. Los precios de las criptomonedas son altamente volátiles, por lo que debe ser consciente de los riesgos potenciales.
- MonedasPrecioCambio en 24H