Best models overall
Ranked on LiveBench general capability and MathArena research math. Pick a model, see why it placed where it did.
| Rank | Model | Combined index | LiveBench | MathArena |
|---|---|---|---|---|
| 1= | Claude Fable 5.1Anthropic | 92.9 | 83.4claude-fable-5-1-max-effort | 87.72% ± 6.03post-releaseClaude-Fable-5.1 (high) |
| 1= | GPT-6 AstraOpenAI | 92.9 | 82.2gpt-6-astra-max | 88.60% ± 5.83post-releaseGPT-6 Astra (max) |
| 3 | Muse Spark 1.3Meta | 57.1 | 81.6muse-spark-1.3-xhigh | 45.61% ± 9.91post-releaseMuse Spark 1.3 |
| 4 | Qwen3.8 2.4T A95BQwen · Open weights | 50.0 | 78.5qwen3.8-max | 57.89% ± 12.82post-releaseQwen3.8-Max |
| 5 | DeepSeek V4.1 FlashDeepSeek · Open weights | 42.9 | 81.1deepseek-v4.1-flash-max | 43.86% ± 9.11post-releaseDeepSeek-V4.1-Flash (Max) |
| 6 | Grok 4.7SpaceXAI | 35.7 | 77.4grok-4.7-xhigh | 46.49% ± 9.78post-releaseGrok 4.7 (xhigh) |
| 7 | Kimi K3MoonshotAI · Open weights | 21.4 | 79.2kimi-k3 | 38.60% ± 12.64Kimi K3 (Think) |
| 8 | Gemini 3.8 FlashGoogle | 7.1 | 75.8gemini-3.8-flash-high | 40.35% ± 9.05post-releaseGemini 3.8 Flash |
51 more with only one source — not ranked
- smaug-flashLiveBench 77.4
- smaug-miniLiveBench 76.9
- union-alphaLiveBench 76.1
- Claude Fable 5LiveBench 83.0
- Claude Opus 4.5LiveBench 72.6
- Claude Opus 4.6LiveBench 74.5
- Claude Opus 4.7LiveBench 76.5
- Claude Opus 4.8LiveBench 76.2
- Claude Opus 5LiveBench 80.1
- Claude Opus 5.5LiveBench 82.1
- Claude Sonnet 4.6LiveBench 73.0
- Claude Sonnet 5LiveBench 76.0
- DeepSeek V4 Flash 0423LiveBench 65.5
- DeepSeek V4 Flash 0731LiveBench 74.2
- DeepSeek V4 Flash Vision ExpLiveBench 76.8
- DeepSeek V4 Pro 0423LiveBench 71.6
- DeepSeek V4 Pro 0813LiveBench 77.4
- Gemini 3.1 Pro PreviewLiveBench 77.0
- Gemini 3.5 FlashLiveBench 74.6
- Gemini 3.5 Flash LiteLiveBench 63.9
- Gemini 3.6 FlashLiveBench 73.6
- Gemini 3.7 FlashLiveBench 78.8
- Muse Spark 1.1LiveBench 75.3
- Muse Spark 1.2LiveBench 78.0
- MiniMax M3LiveBench 67.3
- Kimi K2.6LiveBench 70.5
- Kimi K2.7 CodeLiveBench 68.4
- Nemotron 3 UltraLiveBench 67.4
- GPT-5.2LiveBench 74.6
- GPT-5.2-CodexLiveBench 74.0
- GPT-5.4LiveBench 78.0
- GPT-5.4 MiniLiveBench 66.4
- GPT-5.4 NanoLiveBench 69.6
- GPT-5.5LiveBench 80.2
- GPT-5.6 LunaLiveBench 73.6
- GPT-5.6 SolLiveBench 81.1
- GPT-5.6 TerraLiveBench 77.9
- GPT-6 LunaLiveBench 72.0
- GPT-6 SolLiveBench 79.2
- Qwen3.6 27BLiveBench 64.0
- Qwen3.6 PlusLiveBench 68.9
- Qwen3.7 MaxLiveBench 73.1
- Qwen3.8 27BLiveBench 75.3
- InklingLiveBench 71.9
- Grok 4.3LiveBench 62.2
- Grok 4.5LiveBench 75.8
- Grok 4.6LiveBench 78.0
- Grok Build 0.1LiveBench 67.8
- GLM 5.2LiveBench 73.2
- GLM 5.3LiveBench 76.1
- GLM 5.3 FlashLiveBench 71.6
How the ranking works
- Combined index
- Each model’s LiveBench and MathArena scores are turned into a percentile among the 8 models that have both, then averaged 50/50. Ties use average ranks. It shows relative order, not accuracy.
- Who gets ranked
- Only models with results from both sources. Missing results are never filled with zero. Search and filters don’t change who’s compared.
- Read with care
- Math counts for half, so this isn’t a coding ranking. Use By task for that. Reasoning settings and test dates differ between models. Small gaps aren’t decisive.
- post-release
- MathArena flags models released after its test questions were public. A contamination risk, not proof.
LiveBenchGeneral capability · edition 2026-06-25 · 63 configurations · checked 2026-09-24 · Published results ↗
MathArenaArXivMath · August 2026 · 10 configurations · checked 2026-09-24 · Published results ↗