Model Library / Benchmarks / Terminal-Bench 2.1
Coding & Agents
An agent's ability to complete real-world tasks in a sandboxed command-line terminal. Source: tbench.ai official leaderboard (retrieved Aug 2026).
Bars scaled from 70 so close scores stay legible — the figure at right is the actual accuracy %.