Model Library / Benchmarks / HumanEval
Coding & Agents
Hand-written programming problems scored by unit tests — a classic code-generation benchmark (pass@1).
Bars scaled from 80 so close scores stay legible — the figure at right is the actual pass@1 %.