Model Library / Benchmarks / Video-MME

Multimodal

Video-MME

Multimodal video understanding across short, medium, and long videos in diverse domains. Source: official Video-MME leaderboard (overall, without subtitles).

Features measured

Vision Reasoning

Use cases unlocked

Visual Q&A Chart & diagram analysis

Leaderboard

Higher is better · accuracy %
1 video-SALMONN 2+ 79.7%
2 Gemini 1.5 Pro 75%
3 AdaReTaKe 73.5%
4 InternVL2.5 72.1%
5 JT-VL-Chat 70.6%

Bars scaled from 65 so close scores stay legible — the figure at right is the actual accuracy %.

Benchmark source ↗

← All benchmarks