Model Library / Benchmarks / SWE-bench Verified

Coding & Agents

SWE-bench Verified

Resolves real GitHub issues end-to-end: read the repo, edit code, and pass the hidden tests — an agentic coding benchmark.

Features measured

Code Tool use Reasoning Long context

Use cases unlocked

Autonomous coding agents Bug fixing Coding & code review

Leaderboard

Higher is better · % resolved

Benchmark source ↗

← All benchmarks