Coding & Agents
SWE-bench Multilingual
Resolving real GitHub issues across 300 tasks in 9+ programming languages, scored by whether the patch passes the repo tests. Source: swebench.com multilingual leaderboard (Feb 2026).
Features measured
Code
Tool use
Reasoning
Long context
Multilingual
Use cases unlocked
Autonomous coding agents
Bug fixing
Coding & code review
Leaderboard
Higher is better · % resolved
1
Gemini 3 Flash
72.7%
2
Claude 4.6 Opus
72%
3
Claude 4.5 Opus
70.7%
4
GLM 5
69.7%
5
Gemini 3 Pro
68.7%
6
MiniMax M2.5
68.3%
7
Kimi K2.5
67.3%
8
Claude 4.5 Sonnet
67%
Bars scaled from 60 so close scores stay legible — the figure at right is the actual % resolved.
Benchmark source ↗
← All benchmarks