Model Library / Benchmarks / GDPval

Knowledge & Reasoning

GDPval

Whether AI deliverables across 44 knowledge-work occupations match or beat human experts, graded by expert blind pairwise comparison. Source: GDPval paper (arXiv:2510.04374), gold subset win-or-tie rate.

Features measured

Reasoning Tool use

Use cases unlocked

High-stakes analysis Expert research Drafting & summarising

Leaderboard

Higher is better · win-or-tie %
1 Claude Opus 4.1 47.6%

Benchmark source ↗

← All benchmarks