← All benchmarks

MATH Level 5

Math unit: % independent 88 models scored not in composite

The hardest difficulty band of the MATH dataset: competition problems across algebra, geometry, number theory and precalculus.

What it measures

Reliable symbolic and multi-step arithmetic reasoning.

How to read it

Largely saturated. Useful mostly as a sanity floor for smaller and older models.

Source

Epoch AI Benchmarking Hub

Full ranking

#ModelLabScoreMeasuredMethodSource
1 GPT-5 OpenAI 98.0% 2025-10-29 independent · 2 runs Epoch AI Benchmarking Hub
2 o3 OpenAI 97.8% 2025-04-16 independent Epoch AI Benchmarking Hub
3 o4-mini OpenAI 97.8% 2025-04-16 independent Epoch AI Benchmarking Hub
4 Claude Sonnet 4.5 Anthropic 97.7% 2025-10-21 independent Epoch AI Benchmarking Hub
5 GPT-5 mini OpenAI 97.3% 2025-10-30 independent · 2 runs Epoch AI Benchmarking Hub
6 Qwen3-Max Alibaba 97.1% 2025-10-09 independent Epoch AI Benchmarking Hub
7 DeepSeek-R1 (May 2025)open DeepSeek 96.6% 2025-05-29 independent Epoch AI Benchmarking Hub
8 Gemini 2.5 Pro (May 2025) Google DeepMind 95.9% 2025-05-08 independent Epoch AI Benchmarking Hub
9 o3-mini OpenAI 95.8% 2025-02-13 independent · 2 runs Epoch AI Benchmarking Hub
10 Gemini 2.5 Pro (Mar 2025) Google DeepMind 95.6% 2025-05-07 independent Epoch AI Benchmarking Hub
11 GPT-5 nano OpenAI 95.1% 2025-08-20 independent · 2 runs Epoch AI Benchmarking Hub
12 o1 OpenAI 94.6% 2025-02-13 independent · 2 runs Epoch AI Benchmarking Hub
13 DeepSeek-R1open DeepSeek 93.1% 2025-01-31 independent Epoch AI Benchmarking Hub
14 Claude Haiku 4.5 Anthropic 91.6% 2025-10-22 independent · 2 runs Epoch AI Benchmarking Hub
15 DeepSeek-R1-Distill-Llama-70Bopen DeepSeek 89.9% 2025-03-10 independent Epoch AI Benchmarking Hub
16 Grok-3 mini xAI 89.5% 2025-04-10 independent · 2 runs Epoch AI Benchmarking Hub
17 Grok 3 xAI 88.7% 2025-04-10 independent Epoch AI Benchmarking Hub
18 GPT-4.1 mini OpenAI 87.3% 2025-04-14 independent Epoch AI Benchmarking Hub
19 DeepSeek-R1-Distill-Qwen-14Bopen DeepSeek 87.1% 2025-03-11 independent Epoch AI Benchmarking Hub
20 o1-mini OpenAI 86.7% 2025-02-13 independent · 2 runs Epoch AI Benchmarking Hub
21 Claude Opus 4 Anthropic 85.0% 2025-05-22 independent Epoch AI Benchmarking Hub
22 Claude Sonnet 4 Anthropic 84.4% 2025-05-22 independent Epoch AI Benchmarking Hub
23 Claude 3.7 Sonnet Anthropic 83.9% 2025-03-13 independent · 4 runs Epoch AI Benchmarking Hub
24 Gemini 2.0 Pro Google DeepMind 83.5% 2025-02-06 independent Epoch AI Benchmarking Hub
25 GPT-4.1 OpenAI 83.0% 2025-04-14 independent Epoch AI Benchmarking Hub
26 Gemini 2.0 Flash Google DeepMind 82.2% 2025-02-06 independent Epoch AI Benchmarking Hub
27 Mistral Medium 3 Mistral 81.6% 2025-05-07 independent Epoch AI Benchmarking Hub
28 o1-preview OpenAI 81.6% 2025-01-27 independent Epoch AI Benchmarking Hub
29 GPT-4.5 OpenAI 78.6% 2025-02-28 independent Epoch AI Benchmarking Hub
30 DeepSeek-V3 (Mar 2025)open DeepSeek 75.5% 2025-04-01 independent Epoch AI Benchmarking Hub
31 Gemma 3 27Bopen Google DeepMind 74.0% 2025-03-13 independent Epoch AI Benchmarking Hub
32 Llama 4 Maverickopen Meta AI 73.0% 2025-04-08 independent Epoch AI Benchmarking Hub
33 GPT-4.1 nano OpenAI 70.0% 2025-04-14 independent Epoch AI Benchmarking Hub
34 Qwen3 235B-A22Bopen Alibaba 68.9% 2025-06-03 independent Epoch AI Benchmarking Hub
35 Qwen2.5-Max Alibaba 67.2% 2025-04-01 independent Epoch AI Benchmarking Hub
36 Qwen Plus Alibaba 65.3% 2025-04-07 independent Epoch AI Benchmarking Hub
37 DeepSeek-V3open DeepSeek 64.9% 2025-01-27 independent Epoch AI Benchmarking Hub
38 Phi-4open Microsoft 64.9% 2025-01-31 independent Epoch AI Benchmarking Hub
39 Grok-2 xAI 63.5% 2025-02-17 independent Epoch AI Benchmarking Hub
40 Qwen2.5-72Bopen Alibaba 63.2% 2025-01-27 independent Epoch AI Benchmarking Hub
41 Llama 4 Scoutopen Meta AI 62.3% 2025-04-08 independent Epoch AI Benchmarking Hub
42 Claude 3.5 Sonnet (October 2024) Anthropic 56.9% 2025-01-27 independent Epoch AI Benchmarking Hub
43 Qwen-Turbo Alibaba 56.2% 2025-04-07 independent Epoch AI Benchmarking Hub
44 Qwen2.5-32Bopen Alibaba 56.1% 2025-01-30 independent Epoch AI Benchmarking Hub
45 Gemini 1.5 Pro Google DeepMind 55.6% 2025-01-27 independent · 2 runs Epoch AI Benchmarking Hub
46 GPT-4o mini OpenAI 52.6% 2025-01-27 independent Epoch AI Benchmarking Hub
47 Claude 3.5 Sonnet Anthropic 51.7% 2025-01-27 independent Epoch AI Benchmarking Hub
48 GPT-4o OpenAI 51.4% 2025-02-05 independent · 3 runs Epoch AI Benchmarking Hub
49 Llama 3.1-405Bopen Meta AI 49.8% 2025-01-27 independent Epoch AI Benchmarking Hub
50 Mistral Large 2open Mistral 47.6% 2025-02-25 independent · 2 runs Epoch AI Benchmarking Hub
51 Mistral Small 3.1open Mistral 46.8% 2025-03-18 independent Epoch AI Benchmarking Hub
52 GPT-4 Turbo (Apr 2024) OpenAI 46.7% 2025-02-27 independent Epoch AI Benchmarking Hub
53 Claude 3.5 Haiku Anthropic 46.4% 2025-03-12 independent Epoch AI Benchmarking Hub
54 Mistral Small 3open Mistral 44.8% 2025-01-30 independent Epoch AI Benchmarking Hub
55 Gemini 1.5 Flash Google DeepMind 43.5% 2025-01-27 independent · 2 runs Epoch AI Benchmarking Hub
56 Tulu 3 (Tülu 3) 70Bopen Allen Institute 42.7% 2025-01-27 independent Epoch AI Benchmarking Hub
57 Llama 3.3 70Bopen Meta AI 41.6% 2025-01-27 independent Epoch AI Benchmarking Hub
58 Llama 3.2 90Bopen Meta AI 39.4% 2025-01-27 independent Epoch AI Benchmarking Hub
59 Qwen2-72Bopen Alibaba 39.1% 2025-01-27 independent Epoch AI Benchmarking Hub
60 GPT-4 Turbo (Nov 2023) OpenAI 37.7% 2025-01-27 independent · 2 runs Epoch AI Benchmarking Hub
61 Claude 3 Opus Anthropic 37.5% 2025-01-27 independent Epoch AI Benchmarking Hub
62 Llama 3.1-70Bopen Meta AI 36.7% 2025-01-27 independent Epoch AI Benchmarking Hub
63 Gemma 2 27Bopen Google DeepMind 27.9% 2025-01-27 independent Epoch AI Benchmarking Hub
64 WizardLM-2 8x22Bopen Microsoft 25.7% 2025-01-27 independent Epoch AI Benchmarking Hub
65 Yi-1.5-34Bopen Other 25.5% 2025-01-27 independent Epoch AI Benchmarking Hub
66 Mistral Large Mistral 24.5% 2025-01-27 independent Epoch AI Benchmarking Hub
67 Mixtral 8x22Bopen Mistral 24.2% 2025-01-27 independent Epoch AI Benchmarking Hub
68 GPT-4 (Jun 2023) OpenAI 23.0% 2025-01-27 independent Epoch AI Benchmarking Hub
69 Llama 3.1-8Bopen Meta AI 22.9% 2025-01-27 independent Epoch AI Benchmarking Hub
70 Hermes 2 Theta Llama-3 70Bopen Other 22.7% 2025-01-27 independent Epoch AI Benchmarking Hub
71 Llama 3-70Bopen Meta AI 22.6% 2025-01-27 independent Epoch AI Benchmarking Hub
72 Gemma 2 9Bopen Google DeepMind 21.0% 2025-01-27 independent Epoch AI Benchmarking Hub
73 Claude 3 Sonnet Anthropic 18.2% 2025-01-27 independent Epoch AI Benchmarking Hub
74 phi-3-medium 14Bopen Microsoft 17.6% 2025-01-31 independent Epoch AI Benchmarking Hub
75 Claude 3 Haiku Anthropic 14.9% 2025-01-27 independent Epoch AI Benchmarking Hub
76 Ministral 8Bopen Mistral 14.9% 2025-01-27 independent Epoch AI Benchmarking Hub
77 Ministral 3B Mistral 14.4% 2025-01-27 independent Epoch AI Benchmarking Hub
78 GPT-3.5 Turbo OpenAI 13.8% 2025-01-27 independent · 2 runs Epoch AI Benchmarking Hub
79 Claude 2 Anthropic 11.7% 2025-01-27 independent Epoch AI Benchmarking Hub
80 DBRXopen Other 11.7% 2025-01-27 independent Epoch AI Benchmarking Hub
81 Gemini 1.0 Pro Google DeepMind 11.2% 2025-01-27 independent Epoch AI Benchmarking Hub
82 Mistral NeMoopen Mistral 10.8% 2025-01-27 independent Epoch AI Benchmarking Hub
83 Mixtral 8x7Bopen Mistral 9.6% 2025-01-27 independent · 2 runs Epoch AI Benchmarking Hub
84 DeepSeek LLM 67Bopen DeepSeek 6.4% 2025-01-27 independent Epoch AI Benchmarking Hub
85 Llama 3-8Bopen Meta AI 6.1% 2025-01-27 independent Epoch AI Benchmarking Hub
86 Yi-34Bopen Other 5.1% 2025-01-27 independent Epoch AI Benchmarking Hub
87 Mistral 7Bopen Mistral 3.6% 2025-01-27 independent · 2 runs Epoch AI Benchmarking Hub
88 Llama 2-70Bopen Meta AI 3.3% 2025-01-27 independent Epoch AI Benchmarking Hub