Model Library / Benchmarks / HumanEval

Coding & Agents

HumanEval

Hand-written programming problems scored by unit tests — a classic code-generation benchmark (pass@1).

Features measured

Code

Use cases unlocked

Code generation Developer assistants

Leaderboard

Higher is better · pass@1 %

Bars scaled from 80 so close scores stay legible — the figure at right is the actual pass@1 %.

Benchmark source ↗

← All benchmarks