- AGI0%
- ARC0%
- GPT0%
- AGI+2076.42%
Insightful Beating AI News: Following OpenAI's release of GPT-6 Astra on September 3, various benchmark datasets have been continuously adjusted. Some modifications have enhanced Astra's performance, while scores of certain competing models have declined, raising concerns in the AI community about "leaderboard manipulation" and evaluation transparency.
Notably, Astra's hallucination rate was initially decreased from 4.2% to 2%, then reverted back to 4.2%, whereas GPT-5.6 Sol dropped from 12.2% to 9.4% before returning to 12.2%. In mathematical evaluations, Anthropic's Fable 5.1 score briefly decreased from 87.8% to 78% before recovering to 83%, while GPT-5.6 Sol fluctuated from 83% to 80.5% and back to 83%.
Furthermore, Astra's performance in the ARC-AGI-3 evaluation rose from an initial draft score of 98.6% to a final webpage score of 99.99%, and its programming evaluation score saw a slight increase from 57.7% to 57.9%. OpenAI stated that evaluation results can be influenced by factors such as model versions, tool configurations, reasoning levels, and test runs. The adjustments were made to ensure the data more accurately reflects the model's optimal performance.
However, Stanford University researchers warned that frequent re-evaluation may involve a practice known as "Benchmaxxing," where benchmark scores are maximized by adjusting test conditions. Industry experts highlighted that as AI model competition intensifies, evaluation data has become a crucial tool for measuring model capabilities and market competitiveness. Enhancing the transparency and reproducibility of benchmark testing is increasingly becoming a focal point of attention.
إخلاء المسؤولية: يتم الحصول على المحتوى الحالي من وجهات نظر خارجية أو تتم ترجمته مباشرة بواسطة الذكاء الاصطناعي من وجهات نظر خارجية. لا تضمن CoinEx صحة المحتوى ودقته وأصالته، ولا تشكل أي نصيحة استثمارية من CoinEx. أسعار العملات المشفرة متقلبة للغاية، يرجى الانتباه إلى المخاطر المحتملة.
- العملاتالسعرالتغيرات في 24 ساعة