Model Library / Benchmarks / SWE-bench Verified
Coding & Agents
Resolves real GitHub issues end-to-end: read the repo, edit code, and pass the hidden tests — an agentic coding benchmark.
Benchmark source ↗
← All benchmarks