Original, unpublished research-level mathematics problems commissioned from professional mathematicians. Tiers 1 to 3 span undergraduate through early research difficulty.
Genuine mathematical problem solving on problems that have never appeared in training data.
The problem set is private, so no model can memorise it. Answers are auto-verifiable integers or objects, which keeps grading exact.
| # | Model | Lab | Score | Measured | Method | Source |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 100.0% | 2026-06-10 | independent | Epoch AI Benchmarking Hub |
| 2 | GPT-5.6 Sol | OpenAI | 89.1% | 2026-07-09 | independent | Epoch AI Benchmarking Hub |
| 3 | GPT-5.6 Terra | OpenAI | 86.0% | 2026-07-09 | independent | Epoch AI Benchmarking Hub |
| 4 | Claude Opus 5 | Anthropic | 85.6% | 2026-07-24 | independent | Epoch AI Benchmarking Hub |
| 5 | GPT-5.5 | OpenAI | 85.3% | 2026-06-11 | independent | Epoch AI Benchmarking Hub |
| 6 | GPT-5.4 Pro | OpenAI | 82.5% | 2026-06-13 | independent | Epoch AI Benchmarking Hub |
| 7 | GPT-5.6 Luna | OpenAI | 82.1% | 2026-07-09 | independent | Epoch AI Benchmarking Hub |
| 8 | Claude Opus 4.8 | Anthropic | 80.0% | 2026-06-10 | independent | Epoch AI Benchmarking Hub |
| 9 | Claude Sonnet 4.6 | Anthropic | 80.0% | 2026-02-22 | independent | Epoch AI Benchmarking Hub |
| 10 | GPT-5.4 | OpenAI | 78.6% | 2026-06-11 | independent | Epoch AI Benchmarking Hub |
| 11 | Qwen 3.8 Max | Alibaba | 74.7% | 2026-08-04 | independent | Epoch AI Benchmarking Hub |
| 12 | GPT-5.2 Pro | OpenAI | 74.0% | 2026-06-13 | independent | Epoch AI Benchmarking Hub |
| 13 | Kimi K3open | Moonshot AI | 72.2% | 2026-07-17 | independent | Epoch AI Benchmarking Hub |
| 14 | Gemini 3.7 Flash | Google DeepMind | 71.6% | 2026-08-14 | independent | Epoch AI Benchmarking Hub |
| 15 | Claude Opus 4.7 | Anthropic | 70.2% | 2026-06-10 | independent | Epoch AI Benchmarking Hub |
| 16 | Grok 4.6 | xAI | 66.0% | 2026-08-14 | independent | Epoch AI Benchmarking Hub |
| 17 | Claude Sonnet 5 | Anthropic | 65.6% | 2026-06-30 | independent | Epoch AI Benchmarking Hub |
| 18 | Qwen3.7-Max | Alibaba | 64.6% | 2026-06-13 | independent | Epoch AI Benchmarking Hub |
| 19 | Gemini 3.5 Flash | Google DeepMind | 62.8% | 2026-06-10 | independent | Epoch AI Benchmarking Hub |
| 20 | Gemini 3.1 Pro | Google DeepMind | 59.6% | 2026-06-11 | independent | Epoch AI Benchmarking Hub |
| 21 | GLM-5.2open | Z.ai | 59.2% | 2026-06-19 | independent | Epoch AI Benchmarking Hub |
| 22 | Gemini 3.6 Flash | Google DeepMind | 58.9% | 2026-08-02 | independent | Epoch AI Benchmarking Hub |
| 23 | DeepSeek V4 Flash 0731open | DeepSeek | 57.5% | 2026-08-02 | independent | Epoch AI Benchmarking Hub |
| 24 | Grok 4.5 | xAI | 57.2% | 2026-07-09 | independent | Epoch AI Benchmarking Hub |
| 25 | Kimi K2.6open | Moonshot AI | 57.2% | 2026-06-10 | independent | Epoch AI Benchmarking Hub |
| 26 | GPT-5 Pro | OpenAI | 55.8% | 2026-06-12 | independent | Epoch AI Benchmarking Hub |
| 27 | Kimi K2.7 Codeopen | Moonshot AI | 54.0% | 2026-06-13 | independent | Epoch AI Benchmarking Hub |
| 28 | GPT-5.5 Pro | OpenAI | 51.7% | 2026-04-23 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 29 | Gemini 3 Flash | Google DeepMind | 51.2% | 2026-06-11 | independent | Epoch AI Benchmarking Hub |
| 30 | GPT-5.4 Mini | OpenAI | 51.2% | 2026-06-12 | independent | Epoch AI Benchmarking Hub |
| 31 | Inkling-Smallopen | Thinking Machines | 46.3% | 2026-08-14 | independent | Epoch AI Benchmarking Hub |
| 32 | DeepSeek-V4-Proopen | DeepSeek | 45.3% | 2026-06-17 | independent | Epoch AI Benchmarking Hub |
| 33 | GPT-5.4 Nano | OpenAI | 44.9% | 2026-06-12 | independent | Epoch AI Benchmarking Hub |
| 34 | Grok 4.20 | xAI | 44.9% | 2026-07-13 | independent | Epoch AI Benchmarking Hub |
| 35 | Grok 4.3 Beta | xAI | 42.8% | 2026-06-17 | independent | Epoch AI Benchmarking Hub |
| 36 | Claude Opus 4.6 | Anthropic | 39.7% | 2026-02-12 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 37 | Muse Spark | Meta AI | 39.0% | 2026-04-08 | independent | Epoch AI Benchmarking Hub |
| 38 | Gemini 3 Pro | Google DeepMind | 37.6% | 2025-11-21 | independent | Epoch AI Benchmarking Hub |
| 39 | GPT-5.2 | OpenAI | 36.1% | 2025-12-13 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 40 | GLM-5.1open | Z.ai | 33.4% | 2026-05-11 | independent | Epoch AI Benchmarking Hub |
| 41 | Inklingopen | Thinking Machines | 33.3% | 2026-08-06 | independent | Epoch AI Benchmarking Hub |
| 42 | Kimi K2 Thinkingopen | Moonshot AI | 30.0% | 2025-12-04 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 43 | GPT-5 | OpenAI | 29.8% | 2025-11-13 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 44 | Gemini 2.5 Deep Think | Google DeepMind | 29.0% | — | independent | Epoch AI Benchmarking Hub |
| 45 | Kimi K2.5open | Moonshot AI | 27.9% | 2026-02-03 | independent | Epoch AI Benchmarking Hub |
| 46 | GPT-5.5 Instant | OpenAI | 26.3% | 2026-08-02 | independent | Epoch AI Benchmarking Hub |
| 47 | Qwen 3.6 Plus | Alibaba | 26.2% | 2026-05-12 | independent | Epoch AI Benchmarking Hub |
| 48 | Gemini 3.5 Flash-Lite | Google DeepMind | 26.0% | 2026-08-02 | independent | Epoch AI Benchmarking Hub |
| 49 | GPT-5 mini | OpenAI | 23.8% | 2025-11-13 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 50 | Qwen 3.6 Max (Preview) | Alibaba | 23.1% | 2026-05-27 | independent | Epoch AI Benchmarking Hub |
| 51 | DeepSeek-V3.2open | DeepSeek | 22.1% | 2025-12-22 | independent | Epoch AI Benchmarking Hub |
| 52 | Qwen 3.5 Plus (hosted 397B-A17B) | Alibaba | 21.0% | 2026-05-14 | independent | Epoch AI Benchmarking Hub |
| 53 | Claude Opus 4.5 | Anthropic | 20.6% | 2025-11-25 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 54 | Grok 4 | xAI | 19.7% | 2025-11-13 | independent | Epoch AI Benchmarking Hub |
| 55 | GPT-5.1 | OpenAI | 19.3% | 2025-11-25 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 56 | o4-mini | OpenAI | 18.2% | 2025-11-16 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 57 | GLM-5open | Z.ai | 16.4% | 2026-02-19 | independent | Epoch AI Benchmarking Hub |
| 58 | o3 | OpenAI | 15.1% | 2025-11-17 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 59 | GPT-5 nano | OpenAI | 15.0% | 2025-10-30 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 60 | Claude Sonnet 4.5 | Anthropic | 12.7% | 2025-11-16 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 61 | Gemini 2.5 Pro (Jun 2025) | Google DeepMind | 12.2% | 2025-11-24 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 62 | Qwen 3.6 Flash | Alibaba | 10.3% | 2026-05-12 | independent | Epoch AI Benchmarking Hub |
| 63 | o3-mini | OpenAI | 10.2% | 2025-11-16 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 64 | GPT-4.1 mini | OpenAI | 10.0% | 2025-04-14 | independent | Epoch AI Benchmarking Hub |
| 65 | o1 | OpenAI | 9.3% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 66 | Qwen3-235B-A22B (Jul 2025)open | Alibaba | 8.5% | 2025-12-11 | independent | Epoch AI Benchmarking Hub |
| 67 | Qwen 3.5 Flash (hosted 35B-A3B) | Alibaba | 6.2% | 2026-05-12 | independent | Epoch AI Benchmarking Hub |
| 68 | Claude Haiku 4.5 | Anthropic | 5.0% | 2025-10-22 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 69 | Gemini 2.5 Flash | Google DeepMind | 4.8% | 2025-12-18 | independent | Epoch AI Benchmarking Hub |
| 70 | Claude Opus 4 | Anthropic | 4.3% | 2025-08-05 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 71 | Claude Sonnet 4 | Anthropic | 4.1% | 2025-07-04 | independent | Epoch AI Benchmarking Hub |
| 72 | Grok 3 | xAI | 3.8% | 2025-04-10 | independent | Epoch AI Benchmarking Hub |
| 73 | Claude 3.7 Sonnet | Anthropic | 3.4% | 2025-03-13 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 74 | GLM-4.7open | Z.ai | 2.4% | 2026-01-30 | independent | Epoch AI Benchmarking Hub |
| 75 | DeepSeek-V3open | DeepSeek | 1.7% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 76 | Gemini 2.0 Flash | Google DeepMind | 1.7% | 2025-03-09 | independent | Epoch AI Benchmarking Hub |
| 77 | Qwen2.5-Max | Alibaba | 1.0% | 2025-04-02 | independent | Epoch AI Benchmarking Hub |
| 78 | Grok-2 | xAI | 0.7% | 2025-03-06 | independent | Epoch AI Benchmarking Hub |
| 79 | Llama 4 Maverickopen | Meta AI | 0.7% | 2025-04-08 | independent | Epoch AI Benchmarking Hub |
| 80 | Claude 3.5 Haiku | Anthropic | 0.3% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 81 | GPT-4o | OpenAI | 0.3% | 2025-03-07 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 82 | Claude 3.5 Sonnet | Anthropic | 0.0% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 83 | Claude 3.5 Sonnet (October 2024) | Anthropic | 0.0% | 2025-03-06 | independent | Epoch AI Benchmarking Hub |
| 84 | Claude Opus 4.1 | Anthropic | 0.0% | 2025-08-05 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 85 | DeepSeek-V3 (Mar 2025)open | DeepSeek | 0.0% | 2025-05-08 | independent | Epoch AI Benchmarking Hub |
| 86 | Gemini 1.5 Flash | Google DeepMind | 0.0% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 87 | Gemini 2.0 Pro | Google DeepMind | 0.0% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 88 | GLM-4.5open | Z.ai | 0.0% | 2025-09-08 | independent | Epoch AI Benchmarking Hub |
| 89 | GLM-4.6open | Z.ai | 0.0% | 2025-12-08 | independent | Epoch AI Benchmarking Hub |
| 90 | GPT-4.1 | OpenAI | 0.0% | 2025-04-14 | independent | Epoch AI Benchmarking Hub |
| 91 | GPT-4.1 nano | OpenAI | 0.0% | 2025-04-14 | independent | Epoch AI Benchmarking Hub |
| 92 | Grok-3 mini | xAI | 0.0% | 2025-04-10 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 93 | Llama 4 Scoutopen | Meta AI | 0.0% | 2025-04-08 | independent | Epoch AI Benchmarking Hub |
| 94 | Magistral Small 1.0open | Mistral | 0.0% | 2025-06-10 | independent | Epoch AI Benchmarking Hub |
| 95 | Mistral Large 2open | Mistral | 0.0% | 2025-03-06 | independent | Epoch AI Benchmarking Hub |
| 96 | Mistral Medium 3 | Mistral | 0.0% | 2025-05-07 | independent | Epoch AI Benchmarking Hub |
| 97 | o1-mini | OpenAI | 0.0% | 2025-03-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 98 | Qwen Plus | Alibaba | 0.0% | 2025-05-14 | independent | Epoch AI Benchmarking Hub |
| 99 | Qwen3 235B-A22Bopen | Alibaba | 0.0% | 2025-06-03 | independent | Epoch AI Benchmarking Hub |