- AGI0%
- ARC0%
- GPT0%
- AGI+2076.42%
Insightful Beating AI News: Following OpenAI's release of GPT-6 Astra on September 3, various benchmark datasets have been continuously adjusted. Some modifications have enhanced Astra's performance, while scores of certain competing models have declined, raising concerns in the AI community about "leaderboard manipulation" and evaluation transparency.
Notably, Astra's hallucination rate was initially decreased from 4.2% to 2%, then reverted back to 4.2%, whereas GPT-5.6 Sol dropped from 12.2% to 9.4% before returning to 12.2%. In mathematical evaluations, Anthropic's Fable 5.1 score briefly decreased from 87.8% to 78% before recovering to 83%, while GPT-5.6 Sol fluctuated from 83% to 80.5% and back to 83%.
Furthermore, Astra's performance in the ARC-AGI-3 evaluation rose from an initial draft score of 98.6% to a final webpage score of 99.99%, and its programming evaluation score saw a slight increase from 57.7% to 57.9%. OpenAI stated that evaluation results can be influenced by factors such as model versions, tool configurations, reasoning levels, and test runs. The adjustments were made to ensure the data more accurately reflects the model's optimal performance.
However, Stanford University researchers warned that frequent re-evaluation may involve a practice known as "Benchmaxxing," where benchmark scores are maximized by adjusting test conditions. Industry experts highlighted that as AI model competition intensifies, evaluation data has become a crucial tool for measuring model capabilities and market competitiveness. Enhancing the transparency and reproducibility of benchmark testing is increasingly becoming a focal point of attention.
면책 조항: 현재 콘텐츠는 제3자 관점에서 제공되거나 제3자 관점에서 AI가 직접 번역한 것입니다. CoinEx는 콘텐츠의 진위성, 정확성, 독창성을 보장하지 않으며 CoinEx의 투자 조언으로 간주하지 않습니다. 암호화폐 가격은 변동성이 크므로 잠재적인 위험에 유의하시기 바랍니다.
- 코인가격24시간 변동