- OPUS0%
- HIGHER0%
- GPT0%
Insightful Beating AI News: Perplexity benchmarked GPT-6 Astra with its in-house Agent WANDR, achieving a score of 0.682, the highest among models tested so far. The average cost per task is $11.98. Compared to Claude Fable 5.1, it scored 13.5% higher with a 6.1% lower cost; compared to Opus 5, it scored 27% higher with only a 3.3% higher cost.
WANDR specializes in testing "broad and deep" research tasks. It consists of 500 real research tasks where the Agent is required not only to find a few answers but to identify all objects that meet the criteria, then research each one, verify their identities, and provide each result with a verifiable source. Typical tasks include competitive research, due diligence, literature retrieval, market analysis, and talent search.
Previously, Fable 5.1 ranked first with a score of 0.601 and a cost of $12.76 per task, while Opus 5 scored 0.537 with a cost of $11.60. Astra significantly raised the score this time, and it was not achieved by simply increasing the cost.
면책 조항: 현재 콘텐츠는 제3자 관점에서 제공되거나 제3자 관점에서 AI가 직접 번역한 것입니다. CoinEx는 콘텐츠의 진위성, 정확성, 독창성을 보장하지 않으며 CoinEx의 투자 조언으로 간주하지 않습니다. 암호화폐 가격은 변동성이 크므로 잠재적인 위험에 유의하시기 바랍니다.
- 코인가격24시간 변동