How the models score on public capability benchmarks — and, for each, the features it measures and the real-world use cases those features unlock. Click a benchmark for the full leaderboard.
#1 finishes by model family across all 18 benchmarks. Boards rarely share the same models, so this counts each benchmark's top model rather than a single cross-board score.
Hover a family to see which boards it tops. The full per-benchmark leaderboards are below.
Filter each column from its heading. Type to search — Category also offers a dropdown of values.
| Features measured | Models | |||
|---|---|---|---|---|
| MMLU | Knowledge & Reasoning | Reasoning Multilingual | 5 | Claude Opus 5 89.2% |
| GPQA Diamond | Knowledge & Reasoning | Reasoning | 4 | Claude Opus 5 62.4% |
| SWE-bench Verified | Coding & Agents | Code Tool use Reasoning Long context | 4 | Claude Opus 5 74.5% |
| HumanEval | Coding & Agents | Code | 4 | Claude Opus 5 94.1% |
| MMMU | Multimodal | Vision Reasoning | 3 | Gemini 2.5 Pro 74.3% |
| AIME 2025 | Math | Reasoning | 3 | Claude Opus 5 80% |
| MTEB | Retrieval | Embeddings | 1 | E5 Mistral 7B (embeddings) 66.6 |
| MGSM | Multilingual | Multilingual Reasoning | 3 | Gemini 2.5 Pro 91.5% |
| LMArena (Text Arena) | Knowledge & Reasoning | Reasoning Multilingual | 11 | Claude Mythos 5 1,531 |
| Terminal-Bench 2.1 | Coding & Agents | Tool use Code Reasoning Long context | 8 | Claude Code + Fable 5 83.8% |
| τ²-bench (Telecom) | Coding & Agents | Tool use Reasoning | 4 | Claude 3.7 Sonnet 49% |
| Video-MME | Multimodal | Vision Reasoning | 5 | video-SALMONN 2+ 79.7% |
| SciCode | Coding & Agents | Code Reasoning | 6 | OpenAI o3-mini (low) 10.8% |
| CharXiv (Reasoning) | Multimodal | Vision Reasoning | 7 | Claude 3.5 Sonnet 60.2% |
| MMMLU (Multilingual) | Multilingual | Multilingual Reasoning | 8 | o3 (high) 88.8% |
| GDPval | Knowledge & Reasoning | Reasoning Tool use | 1 | Claude Opus 4.1 47.6% |
| SWE-bench Multilingual | Coding & Agents | Code Tool use Reasoning Long context Multilingual | 8 | Gemini 3 Flash 72.7% |
| OmniDocBench | Multimodal | Vision | 8 | PaddleOCR-VL-1.6 96.3 |
Get an API key and point your first request at ApiSpi in minutes