Independent evidence

Compare models across sources.

A broader view of capability, with every contributing result visible.

Two evaluators · 8 models with both results

LiveBench general capability contributes 50%; MathArena research mathematics contributes 50%. Each score is converted to a rank percentile within the same 8-model group, then averaged. Ties use average ranks. This is a relative index, not an accuracy percentage or a universal quality score.

Both sources are required. Missing evidence stays unranked, never zero-filled. Search and open-weight filters do not change the comparison group. Different reasoning settings and test dates limit comparability; inspect them below. This blend gives substantial weight to mathematics and does not replace a task-specific benchmark.

⚠ MathArena marks models released after the test material became available. This is a contamination risk flag, not proof of contamination. Flags and confidence intervals are retained below; small score differences are not decisive.

LiveBench

General capability · edition 2026-06-25
Checked 2026-09-24 · 63 configurations

Published results ↗

59 results

ModelCombined indexCoverageLiveBench / 100MathArena accuracy
Anthropic: Claude Fable 5.192.9Relative index / 1002 / 2 sources83.4claude-fable-5-1-max-effort87.72% ± 6.03%Claude-Fable-5.1 (high)⚠ Released after test material
OpenAI: GPT-6 Astra92.9Relative index / 1002 / 2 sources82.2gpt-6-astra-max88.60% ± 5.83%GPT-6 Astra (max)⚠ Released after test material
Meta: Muse Spark 1.357.1Relative index / 1002 / 2 sources81.6muse-spark-1.3-xhigh45.61% ± 9.91%Muse Spark 1.3⚠ Released after test material
Qwen: Qwen3.8 2.4T A95B50.0Relative index / 1002 / 2 sources78.5qwen3.8-max57.89% ± 12.82%Qwen3.8-Max⚠ Released after test material
DeepSeek: DeepSeek V4.1 Flash42.9Relative index / 1002 / 2 sources81.1deepseek-v4.1-flash-max43.86% ± 9.11%DeepSeek-V4.1-Flash (Max)⚠ Released after test material
SpaceXAI: Grok 4.735.7Relative index / 1002 / 2 sources77.4grok-4.7-xhigh46.49% ± 9.78%Grok 4.7 (xhigh)⚠ Released after test material
MoonshotAI: Kimi K321.4Relative index / 1002 / 2 sources79.2kimi-k338.60% ± 12.64%Kimi K3 (Think)
Google: Gemini 3.8 Flash7.1Relative index / 1002 / 2 sources75.8gemini-3.8-flash-high40.35% ± 9.05%Gemini 3.8 Flash⚠ Released after test material
smaug-flash—Insufficient coverage1 / 2 sources77.4smaug-flash—
smaug-mini—Insufficient coverage1 / 2 sources76.9smaug-mini—
union-alpha—Insufficient coverage1 / 2 sources76.1union-alpha—
Anthropic: Claude Fable 5—Insufficient coverage1 / 2 sources83.0claude-fable-5-max-effort—
Anthropic: Claude Opus 4.5—Insufficient coverage1 / 2 sources72.6claude-opus-4-5-20251101-thinking-64k-high-effort—
Anthropic: Claude Opus 4.6—Insufficient coverage1 / 2 sources74.5claude-opus-4-6-thinking-auto-high-effort—
Anthropic: Claude Opus 4.7—Insufficient coverage1 / 2 sources76.5claude-opus-4-7-xhigh-effort—
Anthropic: Claude Opus 4.8—Insufficient coverage1 / 2 sources76.2claude-opus-4-8-max-effort—
Anthropic: Claude Opus 5—Insufficient coverage1 / 2 sources80.1claude-opus-5-max-effort—
Anthropic: Claude Opus 5.5—Insufficient coverage1 / 2 sources82.1claude-opus-5-5-xhigh-effort—
Anthropic: Claude Sonnet 4.6—Insufficient coverage1 / 2 sources73.0claude-sonnet-4-6-thinking-auto-medium-effort—
Anthropic: Claude Sonnet 5—Insufficient coverage1 / 2 sources76.0claude-sonnet-5-xhigh-effort—
DeepSeek: DeepSeek V4 Flash 0423—Insufficient coverage1 / 2 sources65.5deepseek-v4-flash—
DeepSeek: DeepSeek V4 Flash 0731—Insufficient coverage1 / 2 sources74.2deepseek-v4-flash-0731—
DeepSeek: DeepSeek V4 Flash Vision Exp—Insufficient coverage1 / 2 sources76.8deepseek-v4-flash-vision-exp—
DeepSeek: DeepSeek V4 Pro 0423—Insufficient coverage1 / 2 sources71.6deepseek-v4-pro—
DeepSeek: DeepSeek V4 Pro 0813—Insufficient coverage1 / 2 sources77.4deepseek-v4-pro-0813—
Google: Gemini 3.1 Pro Preview—Insufficient coverage1 / 2 sources77.0gemini-3.1-pro-preview-high—
Google: Gemini 3.5 Flash—Insufficient coverage1 / 2 sources74.6gemini-3.5-flash-high—
Google: Gemini 3.5 Flash Lite—Insufficient coverage1 / 2 sources63.9gemini-3.5-flash-lite-high—
Google: Gemini 3.6 Flash—Insufficient coverage1 / 2 sources73.6gemini-3.6-flash-high—
Google: Gemini 3.7 Flash—Insufficient coverage1 / 2 sources78.8gemini-3.7-flash-high—
Meta: Muse Spark 1.1—Insufficient coverage1 / 2 sources75.3muse-spark-1.1-xhigh—
Meta: Muse Spark 1.2—Insufficient coverage1 / 2 sources78.0muse-spark-1.2-xhigh—
MiniMax: MiniMax M3—Insufficient coverage1 / 2 sources67.3minimax-m3—
MoonshotAI: Kimi K2.6—Insufficient coverage1 / 2 sources70.5kimi-k2.6-thinking—
MoonshotAI: Kimi K2.7 Code—Insufficient coverage1 / 2 sources68.4kimi-k2.7-code—
NVIDIA: Nemotron 3 Ultra—Insufficient coverage1 / 2 sources67.4nemotron-3-ultra-550b-a55b—
OpenAI: GPT-5.2—Insufficient coverage1 / 2 sources74.6gpt-5.2-2025-12-11-high—
OpenAI: GPT-5.2-Codex—Insufficient coverage1 / 2 sources74.0gpt-5.2-codex—
OpenAI: GPT-5.4—Insufficient coverage1 / 2 sources78.0gpt-5.4-xhigh—
OpenAI: GPT-5.4 Mini—Insufficient coverage1 / 2 sources66.4gpt-5.4-mini-xhigh—
OpenAI: GPT-5.4 Nano—Insufficient coverage1 / 2 sources69.6gpt-5.4-nano-xhigh—
OpenAI: GPT-5.5—Insufficient coverage1 / 2 sources80.2gpt-5.5-xhigh—
OpenAI: GPT-5.6 Luna—Insufficient coverage1 / 2 sources73.6gpt-5.6-luna-max—
OpenAI: GPT-5.6 Sol—Insufficient coverage1 / 2 sources81.1gpt-5.6-sol-max—
OpenAI: GPT-5.6 Terra—Insufficient coverage1 / 2 sources77.9gpt-5.6-terra-max—
OpenAI: GPT-6 Luna—Insufficient coverage1 / 2 sources72.0gpt-6-luna-max—
OpenAI: GPT-6 Sol—Insufficient coverage1 / 2 sources79.2gpt-6-sol-max—
Qwen: Qwen3.6 27B—Insufficient coverage1 / 2 sources64.0qwen3.6-27b—
Qwen: Qwen3.6 Plus—Insufficient coverage1 / 2 sources68.9qwen3.6-plus—
Qwen: Qwen3.7 Max—Insufficient coverage1 / 2 sources73.1qwen3.7-max—
Qwen: Qwen3.8 27B—Insufficient coverage1 / 2 sources75.3qwen3.8-27b—
Thinking Machines: Inkling—Insufficient coverage1 / 2 sources71.9inkling-xhigh—
SpaceXAI: Grok 4.3—Insufficient coverage1 / 2 sources62.2grok-4.3—
SpaceXAI: Grok 4.5—Insufficient coverage1 / 2 sources75.8grok-4.5—
SpaceXAI: Grok 4.6—Insufficient coverage1 / 2 sources78.0grok-4.6—
SpaceXAI: Grok Build 0.1—Insufficient coverage1 / 2 sources67.8grok-build-0.1—
Z.ai: GLM 5.2—Insufficient coverage1 / 2 sources73.2glm-5.2—
Z.ai: GLM 5.3—Insufficient coverage1 / 2 sources76.1glm-5.3—
Z.ai: GLM 5.3 Flash—Insufficient coverage1 / 2 sources71.6glm-5.3-flash—

Coding and data-analysis fit on All models still use their explicitly labeled LiveBench task results. This broader aggregate is separate; mathematical accuracy is not substituted for coding ability.