- GPT0%
- OPUS0%
- GLM0%
Dynamic Beating AI News Flash: Terminal-Bench has released version 4.0, recalibrating the time, CPU, and memory when the agent performs tasks, and fixing 19 tasks. It has removed 8 tasks with saturation, denial, openly disclosed solutions, or quality issues. The maximum execution time for all tasks has been standardized to 8 hours, mainly to reduce interference from timeouts and environmental issues on performance.
In the latest leaderboard, Opus 5 + Claude Code ranks first with 51.8%, followed by Fable 5 at 44.5%. GLM-5.3 + Claude Code scored 41.8%, claiming the third spot, surpassing GPT-5.6 Sol + Codex at 37.3%. Among the top three, GLM-5.3 is the only model not from Anthropic.
In Terminal-Bench 3.0, GLM-5.3 ranked fourth with 32.4%, trailing GPT-5.6 Sol at 34.6%; by version 4.0, GLM-5.3 has climbed to third place, leading Sol by 4.5 percentage points in turn.
Tuyên bố từ chối trách nhiệm: Nội dung hiện tại đến từ ý kiến của bên thứ ba hoặc được AI dịch trực tiếp không đảm bảo tính xác thực, chính xác và độc đáo của nội dung và không cấu thành bất kỳ lời khuyên đầu tư nào liên quan đến CoinEx. Giá tài sản kỹ thuật số biến động dữ dội, vui lòng lưu ý những rủi ro tiềm ẩn.
- Loại coinGiá cảBiên độ 24H