買幣
行情
現貨
合約
理財
活動
更多
reward-center新手專區
信息首頁快訊詳情
OpenAI has been accused of quietly adjusting the GPT-6 Astra evaluation data, with some metrics making competitors "look worse."
  • AGI0%
  • ARC0%
  • GPT0%
  • AGI+2076.42%

Insightful Beating AI News: Following OpenAI's release of GPT-6 Astra on September 3, various benchmark datasets have been continuously adjusted. Some modifications have enhanced Astra's performance, while scores of certain competing models have declined, raising concerns in the AI community about "leaderboard manipulation" and evaluation transparency.

Notably, Astra's hallucination rate was initially decreased from 4.2% to 2%, then reverted back to 4.2%, whereas GPT-5.6 Sol dropped from 12.2% to 9.4% before returning to 12.2%. In mathematical evaluations, Anthropic's Fable 5.1 score briefly decreased from 87.8% to 78% before recovering to 83%, while GPT-5.6 Sol fluctuated from 83% to 80.5% and back to 83%.

Furthermore, Astra's performance in the ARC-AGI-3 evaluation rose from an initial draft score of 98.6% to a final webpage score of 99.99%, and its programming evaluation score saw a slight increase from 57.7% to 57.9%. OpenAI stated that evaluation results can be influenced by factors such as model versions, tool configurations, reasoning levels, and test runs. The adjustments were made to ensure the data more accurately reflects the model's optimal performance.

However, Stanford University researchers warned that frequent re-evaluation may involve a practice known as "Benchmaxxing," where benchmark scores are maximized by adjusting test conditions. Industry experts highlighted that as AI model competition intensifies, evaluation data has become a crucial tool for measuring model capabilities and market competitiveness. Enhancing the transparency and reproducibility of benchmark testing is increasingly becoming a focal point of attention.

來源:BlockBeats

免責聲明:當前內容均來自第三方觀點或由AI直接翻譯第三方觀點,CoinEx不保證內容的真實性、準確性和原創性,不構成CoinEx相關的任何投資建議。數字資產價格波動劇烈,請注意潛在風險。

熱搜榜
  • 幣種
    價格
    24H漲跌