LiveBench
General capability · edition 2026-06-25
Checked 2026-09-24 · 63 configurations
Independent evidence
A broader view of capability, with every contributing result visible.
LiveBench general capability contributes 50%; MathArena research mathematics contributes 50%. Each score is converted to a rank percentile within the same 8-model group, then averaged. Ties use average ranks. This is a relative index, not an accuracy percentage or a universal quality score.
Both sources are required. Missing evidence stays unranked, never zero-filled. Search and open-weight filters do not change the comparison group. Different reasoning settings and test dates limit comparability; inspect them below. This blend gives substantial weight to mathematics and does not replace a task-specific benchmark.
⚠ MathArena marks models released after the test material became available. This is a contamination risk flag, not proof of contamination. Flags and confidence intervals are retained below; small score differences are not decisive.
General capability · edition 2026-06-25
Checked 2026-09-24 · 63 configurations
ArXivMath · August 2026
Checked 2026-09-24 · 10 configurations
| Model | Combined index | Coverage | LiveBench / 100 | MathArena accuracy |
|---|---|---|---|---|
| Anthropic: Claude Fable 5.1 | 92.9Relative index / 100 | 2 / 2 sources | 83.4claude-fable-5-1-max-effort | 87.72% ± 6.03%Claude-Fable-5.1 (high)⚠ Released after test material |
| OpenAI: GPT-6 Astra | 92.9Relative index / 100 | 2 / 2 sources | 82.2gpt-6-astra-max | 88.60% ± 5.83%GPT-6 Astra (max)⚠ Released after test material |
| Meta: Muse Spark 1.3 | 57.1Relative index / 100 | 2 / 2 sources | 81.6muse-spark-1.3-xhigh | 45.61% ± 9.91%Muse Spark 1.3⚠ Released after test material |
| Qwen: Qwen3.8 2.4T A95B | 50.0Relative index / 100 | 2 / 2 sources | 78.5qwen3.8-max | 57.89% ± 12.82%Qwen3.8-Max⚠ Released after test material |
| DeepSeek: DeepSeek V4.1 Flash | 42.9Relative index / 100 | 2 / 2 sources | 81.1deepseek-v4.1-flash-max | 43.86% ± 9.11%DeepSeek-V4.1-Flash (Max)⚠ Released after test material |
| SpaceXAI: Grok 4.7 | 35.7Relative index / 100 | 2 / 2 sources | 77.4grok-4.7-xhigh | 46.49% ± 9.78%Grok 4.7 (xhigh)⚠ Released after test material |
| MoonshotAI: Kimi K3 | 21.4Relative index / 100 | 2 / 2 sources | 79.2kimi-k3 | 38.60% ± 12.64%Kimi K3 (Think) |
| Google: Gemini 3.8 Flash | 7.1Relative index / 100 | 2 / 2 sources | 75.8gemini-3.8-flash-high | 40.35% ± 9.05%Gemini 3.8 Flash⚠ Released after test material |
| smaug-flash | —Insufficient coverage | 1 / 2 sources | 77.4smaug-flash | — |
| smaug-mini | —Insufficient coverage | 1 / 2 sources | 76.9smaug-mini | — |
| union-alpha | —Insufficient coverage | 1 / 2 sources | 76.1union-alpha | — |
| Anthropic: Claude Fable 5 | —Insufficient coverage | 1 / 2 sources | 83.0claude-fable-5-max-effort | — |
| Anthropic: Claude Opus 4.5 | —Insufficient coverage | 1 / 2 sources | 72.6claude-opus-4-5-20251101-thinking-64k-high-effort | — |
| Anthropic: Claude Opus 4.6 | —Insufficient coverage | 1 / 2 sources | 74.5claude-opus-4-6-thinking-auto-high-effort | — |
| Anthropic: Claude Opus 4.7 | —Insufficient coverage | 1 / 2 sources | 76.5claude-opus-4-7-xhigh-effort | — |
| Anthropic: Claude Opus 4.8 | —Insufficient coverage | 1 / 2 sources | 76.2claude-opus-4-8-max-effort | — |
| Anthropic: Claude Opus 5 | —Insufficient coverage | 1 / 2 sources | 80.1claude-opus-5-max-effort | — |
| Anthropic: Claude Opus 5.5 | —Insufficient coverage | 1 / 2 sources | 82.1claude-opus-5-5-xhigh-effort | — |
| Anthropic: Claude Sonnet 4.6 | —Insufficient coverage | 1 / 2 sources | 73.0claude-sonnet-4-6-thinking-auto-medium-effort | — |
| Anthropic: Claude Sonnet 5 | —Insufficient coverage | 1 / 2 sources | 76.0claude-sonnet-5-xhigh-effort | — |
| DeepSeek: DeepSeek V4 Flash 0423 | —Insufficient coverage | 1 / 2 sources | 65.5deepseek-v4-flash | — |
| DeepSeek: DeepSeek V4 Flash 0731 | —Insufficient coverage | 1 / 2 sources | 74.2deepseek-v4-flash-0731 | — |
| DeepSeek: DeepSeek V4 Flash Vision Exp | —Insufficient coverage | 1 / 2 sources | 76.8deepseek-v4-flash-vision-exp | — |
| DeepSeek: DeepSeek V4 Pro 0423 | —Insufficient coverage | 1 / 2 sources | 71.6deepseek-v4-pro | — |
| DeepSeek: DeepSeek V4 Pro 0813 | —Insufficient coverage | 1 / 2 sources | 77.4deepseek-v4-pro-0813 | — |
| Google: Gemini 3.1 Pro Preview | —Insufficient coverage | 1 / 2 sources | 77.0gemini-3.1-pro-preview-high | — |
| Google: Gemini 3.5 Flash | —Insufficient coverage | 1 / 2 sources | 74.6gemini-3.5-flash-high | — |
| Google: Gemini 3.5 Flash Lite | —Insufficient coverage | 1 / 2 sources | 63.9gemini-3.5-flash-lite-high | — |
| Google: Gemini 3.6 Flash | —Insufficient coverage | 1 / 2 sources | 73.6gemini-3.6-flash-high | — |
| Google: Gemini 3.7 Flash | —Insufficient coverage | 1 / 2 sources | 78.8gemini-3.7-flash-high | — |
| Meta: Muse Spark 1.1 | —Insufficient coverage | 1 / 2 sources | 75.3muse-spark-1.1-xhigh | — |
| Meta: Muse Spark 1.2 | —Insufficient coverage | 1 / 2 sources | 78.0muse-spark-1.2-xhigh | — |
| MiniMax: MiniMax M3 | —Insufficient coverage | 1 / 2 sources | 67.3minimax-m3 | — |
| MoonshotAI: Kimi K2.6 | —Insufficient coverage | 1 / 2 sources | 70.5kimi-k2.6-thinking | — |
| MoonshotAI: Kimi K2.7 Code | —Insufficient coverage | 1 / 2 sources | 68.4kimi-k2.7-code | — |
| NVIDIA: Nemotron 3 Ultra | —Insufficient coverage | 1 / 2 sources | 67.4nemotron-3-ultra-550b-a55b | — |
| OpenAI: GPT-5.2 | —Insufficient coverage | 1 / 2 sources | 74.6gpt-5.2-2025-12-11-high | — |
| OpenAI: GPT-5.2-Codex | —Insufficient coverage | 1 / 2 sources | 74.0gpt-5.2-codex | — |
| OpenAI: GPT-5.4 | —Insufficient coverage | 1 / 2 sources | 78.0gpt-5.4-xhigh | — |
| OpenAI: GPT-5.4 Mini | —Insufficient coverage | 1 / 2 sources | 66.4gpt-5.4-mini-xhigh | — |
| OpenAI: GPT-5.4 Nano | —Insufficient coverage | 1 / 2 sources | 69.6gpt-5.4-nano-xhigh | — |
| OpenAI: GPT-5.5 | —Insufficient coverage | 1 / 2 sources | 80.2gpt-5.5-xhigh | — |
| OpenAI: GPT-5.6 Luna | —Insufficient coverage | 1 / 2 sources | 73.6gpt-5.6-luna-max | — |
| OpenAI: GPT-5.6 Sol | —Insufficient coverage | 1 / 2 sources | 81.1gpt-5.6-sol-max | — |
| OpenAI: GPT-5.6 Terra | —Insufficient coverage | 1 / 2 sources | 77.9gpt-5.6-terra-max | — |
| OpenAI: GPT-6 Luna | —Insufficient coverage | 1 / 2 sources | 72.0gpt-6-luna-max | — |
| OpenAI: GPT-6 Sol | —Insufficient coverage | 1 / 2 sources | 79.2gpt-6-sol-max | — |
| Qwen: Qwen3.6 27B | —Insufficient coverage | 1 / 2 sources | 64.0qwen3.6-27b | — |
| Qwen: Qwen3.6 Plus | —Insufficient coverage | 1 / 2 sources | 68.9qwen3.6-plus | — |
| Qwen: Qwen3.7 Max | —Insufficient coverage | 1 / 2 sources | 73.1qwen3.7-max | — |
| Qwen: Qwen3.8 27B | —Insufficient coverage | 1 / 2 sources | 75.3qwen3.8-27b | — |
| Thinking Machines: Inkling | —Insufficient coverage | 1 / 2 sources | 71.9inkling-xhigh | — |
| SpaceXAI: Grok 4.3 | —Insufficient coverage | 1 / 2 sources | 62.2grok-4.3 | — |
| SpaceXAI: Grok 4.5 | —Insufficient coverage | 1 / 2 sources | 75.8grok-4.5 | — |
| SpaceXAI: Grok 4.6 | —Insufficient coverage | 1 / 2 sources | 78.0grok-4.6 | — |
| SpaceXAI: Grok Build 0.1 | —Insufficient coverage | 1 / 2 sources | 67.8grok-build-0.1 | — |
| Z.ai: GLM 5.2 | —Insufficient coverage | 1 / 2 sources | 73.2glm-5.2 | — |
| Z.ai: GLM 5.3 | —Insufficient coverage | 1 / 2 sources | 76.1glm-5.3 | — |
| Z.ai: GLM 5.3 Flash | —Insufficient coverage | 1 / 2 sources | 71.6glm-5.3-flash | — |
Coding and data-analysis fit on All models still use their explicitly labeled LiveBench task results. This broader aggregate is separate; mathematical accuracy is not substituted for coding ability.