- OPUS0%
- HIGHER0%
- GPT0%
Insightful Beating AI News: Perplexity benchmarked GPT-6 Astra with its in-house Agent WANDR, achieving a score of 0.682, the highest among models tested so far. The average cost per task is $11.98. Compared to Claude Fable 5.1, it scored 13.5% higher with a 6.1% lower cost; compared to Opus 5, it scored 27% higher with only a 3.3% higher cost.
WANDR specializes in testing "broad and deep" research tasks. It consists of 500 real research tasks where the Agent is required not only to find a few answers but to identify all objects that meet the criteria, then research each one, verify their identities, and provide each result with a verifiable source. Typical tasks include competitive research, due diligence, literature retrieval, market analysis, and talent search.
Previously, Fable 5.1 ranked first with a score of 0.601 and a cost of $12.76 per task, while Opus 5 scored 0.537 with a cost of $11.60. Astra significantly raised the score this time, and it was not achieved by simply increasing the cost.
免責聲明:當前內容均來自第三方觀點或由AI直接翻譯第三方觀點,CoinEx不保證內容的真實性、準確性和原創性,不構成CoinEx相關的任何投資建議。數字資產價格波動劇烈,請注意潛在風險。
- 幣種價格24H漲跌