暗号資産購入
マーケット
スポット
先物
金融
特別企画
さらに
reward-center新規登録ゾーン
ホーム速報詳細
Let AI migrate the entire codebase, 94.6% did not pass the test: Out of 20 tasks, 13 were left undone
  • OPUS0%

Dynamic Beating AI News: AI Agent company Einsia has released the SWE Refactor Bench, specifically designed to test if a Coding Agent can autonomously refactor an entire codebase. The benchmark consists of 20 tasks from real projects such as SQLite, zlib, libsodium, GraphHopper, involving tasks like translating C to Rust, migrating from Maven to Gradle, and porting SQLite to WASI. Each task was allocated 6 to 30 hours for completion.

The first evaluation stage checks if the old tech stack has indeed been replaced to prevent the Agent from carrying over legacy code. The second stage runs over 130,000 fixed checks. Finally, 6 independent Coding Agents are deployed, each spending 1 hour to specifically search for hidden bugs.

A total of 520 runs were conducted using 8 cutting-edge models and 26 configurations. Out of these, 340 runs successfully completed the migration, with 88 runs passing all pre-defined tests. However, in 60 cases, hidden bugs were discovered in the final round. Only 28 runs successfully passed all stages, resulting in a pass rate of 5.4%. Out of the 20 tasks, 13 were left incomplete. The xhigh configuration of Claude Opus 5 performed the best, completing all 20 tasks.

The current Coding Agents are now capable of large-scale code modifications. However, achieving a state where "the entire system is modified without any errors" is still a distant goal.

ソース:BlockBeats

免責事項:現在のコンテンツは第三者の視点に基づくもの、または第三者の視点からAIが直接翻訳したものです。CoinExはコンテンツの信頼性、正確性、独創性を保証するものではなく、CoinExからの投資アドバイスを構成するものではありません。暗号資産の価格変動は急激に変動します。潜在的なリスクにご注意ください。

検索上位
  • コインリスト
    価格
    24時間価格変動