How to read this page
LiveBench reports category scores on objective tasks. Arena reports blind human-preference Elo. SWE-bench Verified reports the rate at which an agent configuration resolves real repository issues.
Independent benchmark snapshots
Scores remain separated by source and metric. They are evidence to inspect, not a universal model ranking.
Snapshots refresh daily. A failed source never overwrites the last valid data. Exact model versions, test setup, and original leaderboard links are retained for each result.
Retrieved: Jul 19, 2026, 8:44 PM
Arena has no official public API. Data is read from a daily structured mirror and always links back to the official leaderboard.
| Rank | Exact model | Score | Metric | Samples / votes |
|---|---|---|---|---|
| 1 | kimi-k3 | 1679 | Elo (Elo) | 1757 |
| 2 | claude-fable-5 | 1631 | Elo (Elo) | 2505 |
| 3 | gpt-5.6-sol-xhigh (codex-harness) | 1618 | Elo (Elo) | 2542 |
| 4 | glm-5.2 (max) | 1587 | Elo (Elo) | 4722 |
| 5 | claude-opus-4-8-thinking | 1562 | Elo (Elo) | 7309 |
| 6 | grok-4.5 | 1558 | Elo (Elo) | 2214 |
| 7 | claude-opus-4-7-thinking | 1558 | Elo (Elo) | 10534 |
| 8 | claude-opus-4-7 | 1555 | Elo (Elo) | 9976 |
| 9 | claude-opus-4-6-thinking | 1542 | Elo (Elo) | 12919 |
| 10 | claude-sonnet-5-high | 1542 | Elo (Elo) | 2959 |
Arena has no official public API. Data is read from a daily structured mirror and always links back to the official leaderboard.
| Rank | Exact model | Score | Metric | Samples / votes |
|---|---|---|---|---|
| 1 | claude-opus-4-6-search | 1255 | Elo (Elo) | 105728 |
| 2 | gpt-5.5-search | 1239 | Elo (Elo) | 61253 |
| 3 | claude-opus-4-7 | 1235 | Elo (Elo) | 62150 |
| 4 | claude-fable-5 | 1230 | Elo (Elo) | 9895 |
| 5 | ernie-5.1 | 1227 | Elo (Elo) | 3803 |
| 6 | claude-sonnet-4-6-search | 1222 | Elo (Elo) | 105370 |
| 7 | gemini-3.1-pro-grounding | 1211 | Elo (Elo) | 83445 |
| 8 | claude-opus-4-8 | 1207 | Elo (Elo) | 42248 |
| 9 | gemini-3-pro-grounding | 1207 | Elo (Elo) | 37277 |
| 10 | gpt-5.2-search | 1206 | Elo (Elo) | 52708 |
LiveBench reports category scores on objective tasks. Arena reports blind human-preference Elo. SWE-bench Verified reports the rate at which an agent configuration resolves real repository issues.
Scores from different benchmarks cannot be added together. The everyday-life view uses a preference proxy because no general authoritative benchmark covers that scenario.