Model Library / Benchmarks / Terminal-Bench 2.1

Coding & Agents

Terminal-Bench 2.1

An agent's ability to complete real-world tasks in a sandboxed command-line terminal. Source: tbench.ai official leaderboard (retrieved Aug 2026).

Features measured

Tool use Code Reasoning Long context

Use cases unlocked

Autonomous coding agents Tool use & function calling Bug fixing

Leaderboard

Higher is better · accuracy %
1 Claude Code + Fable 5 83.8%
2 Codex + GPT-5.5 83.1%
3 Terminus 2 + Fable 5 80.4%
4 Cursor CLI + Grok 4.5 79.3%
5 Claude Code + Opus 4.8 78.9%
6 Codex + GPT-5.6 Terra 78.4%
7 Terminus 2 + GPT-5.5 78%
8 mini-SWE-agent + Muse Spark 1.1 76.2%

Bars scaled from 70 so close scores stay legible — the figure at right is the actual accuracy %.

Benchmark source ↗

← All benchmarks