- GPT0%
- OPUS0%
- GLM0%
Dynamic Beating AI News Flash: Terminal-Bench has released version 4.0, recalibrating the time, CPU, and memory when the agent performs tasks, and fixing 19 tasks. It has removed 8 tasks with saturation, denial, openly disclosed solutions, or quality issues. The maximum execution time for all tasks has been standardized to 8 hours, mainly to reduce interference from timeouts and environmental issues on performance.
In the latest leaderboard, Opus 5 + Claude Code ranks first with 51.8%, followed by Fable 5 at 44.5%. GLM-5.3 + Claude Code scored 41.8%, claiming the third spot, surpassing GPT-5.6 Sol + Codex at 37.3%. Among the top three, GLM-5.3 is the only model not from Anthropic.
In Terminal-Bench 3.0, GLM-5.3 ranked fourth with 32.4%, trailing GPT-5.6 Sol at 34.6%; by version 4.0, GLM-5.3 has climbed to third place, leading Sol by 4.5 percentage points in turn.
Disclaimer: The current content is sourced from third-party perspectives or directly translated by AI from third-party perspectives. CoinEx does not guarantee the authenticity, accuracy, and originality of the content, and it does not constitute any investment advice from CoinEx. The prices of cryptocurrencies are highly volatile, please be aware of the potential risks.
- CoinsPrice24H Change