코인 매입
시장
현물
선물
재테크
이벤트
더 알아보기
reward-center초보자 존
홈 피드빠른 소식 정보
Let AI migrate the entire codebase, 94.6% did not pass the test: Out of 20 tasks, 13 were left undone
  • OPUS0%

Dynamic Beating AI News: AI Agent company Einsia has released the SWE Refactor Bench, specifically designed to test if a Coding Agent can autonomously refactor an entire codebase. The benchmark consists of 20 tasks from real projects such as SQLite, zlib, libsodium, GraphHopper, involving tasks like translating C to Rust, migrating from Maven to Gradle, and porting SQLite to WASI. Each task was allocated 6 to 30 hours for completion.

The first evaluation stage checks if the old tech stack has indeed been replaced to prevent the Agent from carrying over legacy code. The second stage runs over 130,000 fixed checks. Finally, 6 independent Coding Agents are deployed, each spending 1 hour to specifically search for hidden bugs.

A total of 520 runs were conducted using 8 cutting-edge models and 26 configurations. Out of these, 340 runs successfully completed the migration, with 88 runs passing all pre-defined tests. However, in 60 cases, hidden bugs were discovered in the final round. Only 28 runs successfully passed all stages, resulting in a pass rate of 5.4%. Out of the 20 tasks, 13 were left incomplete. The xhigh configuration of Claude Opus 5 performed the best, completing all 20 tasks.

The current Coding Agents are now capable of large-scale code modifications. However, achieving a state where "the entire system is modified without any errors" is still a distant goal.

출처:BlockBeats

면책 조항: 현재 콘텐츠는 제3자 관점에서 제공되거나 제3자 관점에서 AI가 직접 번역한 것입니다. CoinEx는 콘텐츠의 진위성, 정확성, 독창성을 보장하지 않으며 CoinEx의 투자 조언으로 간주하지 않습니다. 암호화폐 가격은 변동성이 크므로 잠재적인 위험에 유의하시기 바랍니다.

인기 검색
  • 코인
    가격
    24시간 변동