← All benchmarks

FrontierMath Tier 4

Math unit: % independent 82 models scored counts toward composite (weight 1.5)

The hardest tier of FrontierMath: problems that would occupy a specialist researcher for days. Designed to stay unsaturated for years.

What it measures

The absolute frontier of automated mathematical reasoning.

How to read it

Small problem count, so a single solved problem moves the percentage a lot. Treat small gaps between models as noise.

Source

Epoch AI Benchmarking Hub

Full ranking

#ModelLabScoreMeasuredMethodSource
1 Claude Fable 5 Anthropic 100.0% 2026-06-10 independent Epoch AI Benchmarking Hub
2 GPT-5.6 Sol OpenAI 81.7% 2026-07-09 independent · 2 runs Epoch AI Benchmarking Hub
3 AI Co-Mathematician Google DeepMind 75.6% 2026-06-12 independent Epoch AI Benchmarking Hub
4 Claude Opus 5 Anthropic 73.2% 2026-07-24 independent Epoch AI Benchmarking Hub
5 GPT-5.5 OpenAI 72.5% 2026-06-11 independent Epoch AI Benchmarking Hub
6 GPT-5.6 Terra OpenAI 70.7% 2026-07-09 independent Epoch AI Benchmarking Hub
7 GPT-5.6 Luna OpenAI 61.0% 2026-07-09 independent Epoch AI Benchmarking Hub
8 GPT-5.4 Pro OpenAI 58.5% 2026-06-13 independent Epoch AI Benchmarking Hub
9 Claude Opus 4.8 Anthropic 56.1% 2026-06-10 independent Epoch AI Benchmarking Hub
10 GPT-5.4 OpenAI 50.0% 2026-04-02 independent · 2 runs Epoch AI Benchmarking Hub
11 Qwen 3.8 Max Alibaba 46.3% 2026-08-04 independent Epoch AI Benchmarking Hub
12 GPT-5.2 Pro OpenAI 46.0% 2026-06-13 independent Epoch AI Benchmarking Hub
13 GPT-5.5 Pro OpenAI 39.6% 2026-04-23 independent · 2 runs Epoch AI Benchmarking Hub
14 Kimi K3open Moonshot AI 39.0% 2026-07-17 independent Epoch AI Benchmarking Hub
15 Gemini 3.7 Flash Google DeepMind 36.6% 2026-08-14 independent Epoch AI Benchmarking Hub
16 Qwen3.7-Max Alibaba 34.1% 2026-06-13 independent Epoch AI Benchmarking Hub
17 Claude Opus 4.7 Anthropic 31.7% 2026-06-10 independent Epoch AI Benchmarking Hub
18 Grok 4.6 xAI 31.7% 2026-08-14 independent Epoch AI Benchmarking Hub
19 Claude Sonnet 5 Anthropic 29.3% 2026-06-30 independent Epoch AI Benchmarking Hub
20 GLM-5.2open Z.ai 29.3% 2026-06-19 independent Epoch AI Benchmarking Hub
21 Gemini 3.1 Pro Google DeepMind 26.8% 2026-06-11 independent Epoch AI Benchmarking Hub
22 Gemini 3.5 Flash Google DeepMind 26.8% 2026-06-10 independent Epoch AI Benchmarking Hub
23 Kimi K2.6open Moonshot AI 25.6% 2026-06-10 independent Epoch AI Benchmarking Hub
24 DeepSeek V4 Flash 0731open DeepSeek 24.4% 2026-08-02 independent Epoch AI Benchmarking Hub
25 Grok 4.5 xAI 24.4% 2026-07-09 independent Epoch AI Benchmarking Hub
26 Gemini 3.6 Flash Google DeepMind 22.0% 2026-08-02 independent Epoch AI Benchmarking Hub
27 Claude Opus 4.6 Anthropic 19.8% 2026-02-12 independent · 4 runs Epoch AI Benchmarking Hub
28 GPT-5 Pro OpenAI 19.5% 2026-06-12 independent Epoch AI Benchmarking Hub
29 Gemini 3 Pro Google DeepMind 18.8% 2025-11-21 independent Epoch AI Benchmarking Hub
30 Gemini 3 Flash Google DeepMind 17.1% 2026-06-11 independent Epoch AI Benchmarking Hub
31 Grok 4.20 xAI 17.1% 2026-07-13 independent Epoch AI Benchmarking Hub
32 Inkling-Smallopen Thinking Machines 17.1% 2026-08-14 independent Epoch AI Benchmarking Hub
33 GPT-5.2 OpenAI 15.1% 2025-12-14 independent · 4 runs Epoch AI Benchmarking Hub
34 Grok 4.3 Beta xAI 14.6% 2026-06-17 independent Epoch AI Benchmarking Hub
35 Muse Spark Meta AI 14.6% 2026-04-08 independent Epoch AI Benchmarking Hub
36 GPT-5.4 Nano OpenAI 12.2% 2026-06-12 independent Epoch AI Benchmarking Hub
37 Kimi K2.7 Codeopen Moonshot AI 12.2% 2026-06-13 independent Epoch AI Benchmarking Hub
38 Gemini 2.5 Deep Think Google DeepMind 10.4% independent Epoch AI Benchmarking Hub
39 GPT-5.4 Mini OpenAI 9.8% 2026-06-12 independent Epoch AI Benchmarking Hub
40 GPT-5.1 OpenAI 8.3% 2025-11-17 independent · 2 runs Epoch AI Benchmarking Hub
41 Inklingopen Thinking Machines 4.9% 2026-08-06 independent Epoch AI Benchmarking Hub
42 Gemini 2.5 Flash Google DeepMind 4.2% 2025-12-18 independent Epoch AI Benchmarking Hub
43 Kimi K2.5open Moonshot AI 4.2% 2026-02-02 independent Epoch AI Benchmarking Hub
44 Qwen 3.6 Max (Preview) Alibaba 4.2% 2026-05-28 independent Epoch AI Benchmarking Hub
45 Claude Opus 4.5 Anthropic 3.5% 2025-11-25 independent · 3 runs Epoch AI Benchmarking Hub
46 Claude Sonnet 4.5 Anthropic 3.1% 2025-10-22 independent · 2 runs Epoch AI Benchmarking Hub
47 Claude Opus 4.1 Anthropic 2.4% 2026-06-11 independent Epoch AI Benchmarking Hub
48 DeepSeek-V4-Proopen DeepSeek 2.4% 2026-06-17 independent Epoch AI Benchmarking Hub
49 GPT-5.5 Instant OpenAI 2.4% 2026-08-02 independent Epoch AI Benchmarking Hub
50 Claude Haiku 4.5 Anthropic 2.1% 2025-10-22 independent Epoch AI Benchmarking Hub
51 Claude Opus 4 Anthropic 2.1% 2025-07-01 independent · 2 runs Epoch AI Benchmarking Hub
52 DeepSeek-V3.2open DeepSeek 2.1% 2025-12-16 independent Epoch AI Benchmarking Hub
53 GLM-4.6open Z.ai 2.1% 2025-12-08 independent Epoch AI Benchmarking Hub
54 GLM-5open Z.ai 2.1% 2026-02-19 independent Epoch AI Benchmarking Hub
55 Grok 4 Heavy xAI 2.1% 2025-10-09 independent Epoch AI Benchmarking Hub
56 o3 OpenAI 2.1% 2025-07-01 independent Epoch AI Benchmarking Hub
57 Qwen 3.5 Plus (hosted 397B-A17B) Alibaba 2.1% 2026-05-15 independent Epoch AI Benchmarking Hub
58 Claude 3.5 Sonnet Anthropic 0.0% 2025-07-01 independent Epoch AI Benchmarking Hub
59 Claude 3.5 Sonnet (October 2024) Anthropic 0.0% 2025-07-01 independent Epoch AI Benchmarking Hub
60 Claude 3.7 Sonnet Anthropic 0.0% 2025-07-01 independent Epoch AI Benchmarking Hub
61 Claude Sonnet 4 Anthropic 0.0% 2025-07-01 independent · 2 runs Epoch AI Benchmarking Hub
62 Claude Sonnet 4.6 Anthropic 0.0% 2026-02-22 independent Epoch AI Benchmarking Hub
63 DeepSeek-R1 (May 2025)open DeepSeek 0.0% 2025-07-01 independent Epoch AI Benchmarking Hub
64 Gemini 2.5 Pro (Jun 2025) Google DeepMind 0.0% 2025-07-03 independent · 2 runs Epoch AI Benchmarking Hub
65 Gemini 3.5 Flash-Lite Google DeepMind 0.0% 2026-08-02 independent Epoch AI Benchmarking Hub
66 GLM-4.5open Z.ai 0.0% 2025-09-08 independent Epoch AI Benchmarking Hub
67 GLM-4.7open Z.ai 0.0% 2026-01-30 independent Epoch AI Benchmarking Hub
68 GLM-5.1open Z.ai 0.0% 2026-05-12 independent Epoch AI Benchmarking Hub
69 GPT-4.1 OpenAI 0.0% 2025-07-01 independent Epoch AI Benchmarking Hub
70 GPT-5 OpenAI 0.0% 2025-10-30 independent · 2 runs Epoch AI Benchmarking Hub
71 GPT-5 mini OpenAI 0.0% 2025-10-30 independent · 2 runs Epoch AI Benchmarking Hub
72 GPT-5 nano OpenAI 0.0% 2025-10-30 independent · 2 runs Epoch AI Benchmarking Hub
73 Grok 3 xAI 0.0% 2025-07-01 independent Epoch AI Benchmarking Hub
74 Grok 4 xAI 0.0% 2025-08-11 independent Epoch AI Benchmarking Hub
75 Grok-3 mini xAI 0.0% 2025-07-01 independent Epoch AI Benchmarking Hub
76 Kimi K2 Thinkingopen Moonshot AI 0.0% 2025-12-08 independent Epoch AI Benchmarking Hub
77 o3-mini OpenAI 0.0% 2026-06-11 independent Epoch AI Benchmarking Hub
78 o4-mini OpenAI 0.0% 2025-08-07 independent · 2 runs Epoch AI Benchmarking Hub
79 Qwen 3.5 Flash (hosted 35B-A3B) Alibaba 0.0% 2026-05-12 independent Epoch AI Benchmarking Hub
80 Qwen 3.6 Flash Alibaba 0.0% 2026-05-12 independent Epoch AI Benchmarking Hub
81 Qwen 3.6 Plus Alibaba 0.0% 2026-05-12 independent Epoch AI Benchmarking Hub
82 Qwen3-235B-A22B (Jul 2025)open Alibaba 0.0% 2025-12-11 independent Epoch AI Benchmarking Hub