Best models overall

Ranked on LiveBench general capability and MathArena research math. Pick a model, see why it placed where it did.

Updated 2026-09-24
8 ranked · 59 tracked

RankModelCombined indexLiveBenchMathArena
1=Claude Fable 5.1Anthropic
92.9
83.4claude-fable-5-1-max-effort87.72% ± 6.03post-releaseClaude-Fable-5.1 (high)
1=GPT-6 AstraOpenAI
92.9
82.2gpt-6-astra-max88.60% ± 5.83post-releaseGPT-6 Astra (max)
3Muse Spark 1.3Meta
57.1
81.6muse-spark-1.3-xhigh45.61% ± 9.91post-releaseMuse Spark 1.3
4Qwen3.8 2.4T A95BQwen · Open weights
50.0
78.5qwen3.8-max57.89% ± 12.82post-releaseQwen3.8-Max
5DeepSeek V4.1 FlashDeepSeek · Open weights
42.9
81.1deepseek-v4.1-flash-max43.86% ± 9.11post-releaseDeepSeek-V4.1-Flash (Max)
6Grok 4.7SpaceXAI
35.7
77.4grok-4.7-xhigh46.49% ± 9.78post-releaseGrok 4.7 (xhigh)
7Kimi K3MoonshotAI · Open weights
21.4
79.2kimi-k338.60% ± 12.64Kimi K3 (Think)
8Gemini 3.8 FlashGoogle
7.1
75.8gemini-3.8-flash-high40.35% ± 9.05post-releaseGemini 3.8 Flash
51 more with only one source — not ranked

How the ranking works

Combined index
Each model’s LiveBench and MathArena scores are turned into a percentile among the 8 models that have both, then averaged 50/50. Ties use average ranks. It shows relative order, not accuracy.
Who gets ranked
Only models with results from both sources. Missing results are never filled with zero. Search and filters don’t change who’s compared.
Read with care
Math counts for half, so this isn’t a coding ranking. Use By task for that. Reasoning settings and test dates differ between models. Small gaps aren’t decisive.
post-release
MathArena flags models released after its test questions were public. A contamination risk, not proof.

LiveBenchGeneral capability · edition 2026-06-25 · 63 configurations · checked 2026-09-24 · Published results ↗

MathArenaArXivMath · August 2026 · 10 configurations · checked 2026-09-24 · Published results ↗