Model Library / Benchmarks / CharXiv (Reasoning)

Multimodal

CharXiv (Reasoning)

Open-ended reasoning over real, complex scientific charts from arXiv papers. Source: CharXiv leaderboard v1.0 (reasoning split, updated Dec 2024).

Features measured

Vision Reasoning

Use cases unlocked

Chart & diagram analysis Document & image understanding Visual Q&A

Leaderboard

Higher is better · accuracy %
1 Claude 3.5 Sonnet 60.2%
2 GPT-4o 47.1%
3 Gemini 1.5 Pro 43.3%
4 InternVL Chat V2.0 Pro 39.8%
5 GPT-4V 37.1%
6 GPT-4o mini 34.1%
7 Gemini 1.5 Flash 33.9%

Benchmark source ↗

← All benchmarks