Model Library / Benchmarks / MMLU

Knowledge & Reasoning

MMLU

Massive Multitask Language Understanding — 57 subjects spanning STEM, humanities, and social sciences.

Features measured

Reasoning Multilingual

Use cases unlocked

Knowledge Q&A General assistants Exam-style reasoning

Leaderboard

Higher is better · accuracy %

Bars scaled from 75 so close scores stay legible — the figure at right is the actual accuracy %.

Benchmark source ↗

← All benchmarks