- OPUS0%
- HIGHER0%
- GPT0%
Insightful Beating AI News: Perplexity benchmarked GPT-6 Astra with its in-house Agent WANDR, achieving a score of 0.682, the highest among models tested so far. The average cost per task is $11.98. Compared to Claude Fable 5.1, it scored 13.5% higher with a 6.1% lower cost; compared to Opus 5, it scored 27% higher with only a 3.3% higher cost.
WANDR specializes in testing "broad and deep" research tasks. It consists of 500 real research tasks where the Agent is required not only to find a few answers but to identify all objects that meet the criteria, then research each one, verify their identities, and provide each result with a verifiable source. Typical tasks include competitive research, due diligence, literature retrieval, market analysis, and talent search.
Previously, Fable 5.1 ranked first with a score of 0.601 and a cost of $12.76 per task, while Opus 5 scored 0.537 with a cost of $11.60. Astra significantly raised the score this time, and it was not achieved by simply increasing the cost.
Tuyên bố từ chối trách nhiệm: Nội dung hiện tại đến từ ý kiến của bên thứ ba hoặc được AI dịch trực tiếp không đảm bảo tính xác thực, chính xác và độc đáo của nội dung và không cấu thành bất kỳ lời khuyên đầu tư nào liên quan đến CoinEx. Giá tài sản kỹ thuật số biến động dữ dội, vui lòng lưu ý những rủi ro tiềm ẩn.
- Loại coinGiá cảBiên độ 24H