448 graduate-level multiple-choice questions in biology, physics and chemistry, written so that PhD holders in the field score around 65% while skilled non-experts with web access stay near 34%.
Deep domain knowledge and multi-step scientific reasoning that cannot be answered by search alone.
The Diamond subset is the hardest slice of GPQA. Scores are run independently by Epoch AI, so they can sit below the numbers a lab reports for the same model.
| # | Model | Lab | Score | Measured | Method | Source |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.7 Flash | Google DeepMind | 94.8% | 2026-08-14 | independent | Epoch AI Benchmarking Hub |
| 2 | GPT-5.4 Pro | OpenAI | 94.6% | 2026-03-20 | independent | Epoch AI Benchmarking Hub |
| 3 | Gemini 3.1 Pro | Google DeepMind | 94.3% | 2026-08-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 4 | GPT-5.5 Pro | OpenAI | 93.9% | 2026-04-24 | independent | Epoch AI Benchmarking Hub |
| 5 | Grok 4.6 | xAI | 93.6% | 2026-08-14 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 6 | Grok 4.5 | xAI | 93.4% | 2026-07-08 | independent | Epoch AI Benchmarking Hub |
| 7 | Qwen 3.8 Max | Alibaba | 92.7% | 2026-08-04 | independent | Epoch AI Benchmarking Hub |
| 8 | Gemini 3 Pro | Google DeepMind | 92.6% | 2025-11-19 | independent | Epoch AI Benchmarking Hub |
| 9 | Claude Opus 5 | Anthropic | 91.6% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 10 | DeepSeek V4 Flash 0731open | DeepSeek | 91.0% | 2026-08-02 | independent | Epoch AI Benchmarking Hub |
| 11 | MiniMax-M3open | MiniMax | 90.9% | 2026-08-10 | independent | Epoch AI Benchmarking Hub |
| 12 | Qwen3.7-Max | Alibaba | 90.9% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 13 | Kimi K2.6open | Moonshot AI | 90.8% | 2026-05-01 | independent | Epoch AI Benchmarking Hub |
| 14 | Kimi K3open | Moonshot AI | 90.0% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 15 | GLM-5.1open | Z.ai | 89.9% | 2026-08-10 | independent | Epoch AI Benchmarking Hub |
| 16 | Muse Spark | Meta AI | 89.8% | 2026-04-08 | independent | Epoch AI Benchmarking Hub |
| 17 | Gemini 3.5 Flash | Google DeepMind | 89.4% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 18 | Grok 4.20 | xAI | 89.3% | 2026-07-13 | independent | Epoch AI Benchmarking Hub |
| 19 | Claude Opus 4.6 | Anthropic | 89.2% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 20 | Gemini 3.6 Flash | Google DeepMind | 88.8% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 21 | Grok 4.3 Beta | xAI | 88.8% | 2026-06-17 | independent | Epoch AI Benchmarking Hub |
| 22 | GPT-5.6 Sol | OpenAI | 88.7% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 23 | Inkling-Smallopen | Thinking Machines | 88.5% | 2026-08-14 | independent | Epoch AI Benchmarking Hub |
| 24 | Qwen 3.6 Plus | Alibaba | 88.4% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 25 | Claude Opus 4.7 | Anthropic | 88.3% | 2026-08-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 26 | Claude Opus 4.8 | Anthropic | 88.3% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 27 | Inklingopen | Thinking Machines | 88.3% | 2026-08-05 | independent | Epoch AI Benchmarking Hub |
| 28 | Kimi K2.7 Codeopen | Moonshot AI | 87.9% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 29 | GLM-5open | Z.ai | 87.8% | 2026-02-12 | independent | Epoch AI Benchmarking Hub |
| 30 | Kimi K2.5open | Moonshot AI | 87.6% | 2026-02-02 | independent | Epoch AI Benchmarking Hub |
| 31 | Qwen 3.6 Max (Preview) | Alibaba | 87.4% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 32 | GPT-5.5 | OpenAI | 87.3% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 33 | Grok 4 | xAI | 87.0% | — | independent | Epoch AI Benchmarking Hub |
| 34 | Gemini 3 Flash | Google DeepMind | 86.3% | 2026-08-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 35 | GPT-5.4 | OpenAI | 86.3% | 2026-07-15 | independent · 5 runs | Epoch AI Benchmarking Hub |
| 36 | Qwen3.5 397B-A17Bopen | Alibaba | 86.1% | 2026-08-07 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 37 | GPT-5.6 Terra | OpenAI | 86.0% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 38 | Claude Sonnet 5 | Anthropic | 85.4% | 2026-08-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 39 | Qwen3.6 27Bopen | Alibaba | 85.4% | 2026-08-07 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 40 | Gemini 2.5 Pro (Jun 2025) | Google DeepMind | 85.1% | 2025-11-16 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 41 | Qwen 3.5 Plus (hosted 397B-A17B) | Alibaba | 84.8% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 42 | Qwen3.7-Plus | Alibaba | 84.8% | 2026-08-07 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 43 | GPT-5.2 | OpenAI | 84.7% | 2026-07-13 | independent · 5 runs | Epoch AI Benchmarking Hub |
| 44 | DeepSeek-V4-Proopen | DeepSeek | 84.6% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 45 | Qwen 3.6 35B-A3Bopen | Alibaba | 84.3% | 2026-08-07 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 46 | Kimi K2 Thinkingopen | Moonshot AI | 84.2% | 2025-11-11 | independent | Epoch AI Benchmarking Hub |
| 47 | Claude Opus 4.5 | Anthropic | 84.1% | 2025-11-25 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 48 | Gemini 2.5 Pro (Mar 2025) | Google DeepMind | 83.8% | 2025-03-31 | independent | Epoch AI Benchmarking Hub |
| 49 | GLM-4.7open | Z.ai | 83.3% | 2026-01-29 | independent | Epoch AI Benchmarking Hub |
| 50 | Qwen 3.6 Flash | Alibaba | 83.3% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 51 | Claude Sonnet 4.6 | Anthropic | 83.2% | 2026-08-06 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 52 | Claude Fable 5 | Anthropic | 82.7% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 53 | GPT-5.5 Instant | OpenAI | 82.5% | 2026-08-02 | independent | Epoch AI Benchmarking Hub |
| 54 | Qwen 3.5 Flash (hosted 35B-A3B) | Alibaba | 82.3% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 55 | Qwen3.7 Flash | Alibaba | 81.6% | 2026-08-07 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 56 | GLM-5.2open | Z.ai | 81.5% | 2026-08-10 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 57 | GPT-5 | OpenAI | 81.1% | 2026-07-20 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 58 | o3 | OpenAI | 80.8% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 59 | Qwen3-235B-A22B (Jul 2025)open | Alibaba | 80.1% | 2025-12-11 | independent | Epoch AI Benchmarking Hub |
| 60 | GPT-5.1 | OpenAI | 79.8% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 61 | GPT-5.6 Luna | OpenAI | 79.2% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 62 | Claude Sonnet 4.5 | Anthropic | 79.1% | 2025-10-28 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 63 | GPT-5.4 Mini | OpenAI | 78.2% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 64 | Gemini 3.5 Flash-Lite | Google DeepMind | 77.8% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 65 | o4-mini | OpenAI | 77.5% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 66 | DeepSeek-V3.2open | DeepSeek | 77.3% | 2026-07-16 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 67 | Gemini 3.1 Flash-Lite | Google DeepMind | 76.6% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 68 | DeepSeek-R1 (May 2025)open | DeepSeek | 76.3% | 2025-05-29 | independent | Epoch AI Benchmarking Hub |
| 69 | Claude Opus 4.1 | Anthropic | 75.8% | 2025-08-05 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 70 | Gemma 4 31B ITopen | Google DeepMind | 75.8% | 2026-08-06 | independent | Epoch AI Benchmarking Hub |
| 71 | gpt-oss-120bopen | OpenAI | 75.8% | 2025-12-11 | independent | Epoch AI Benchmarking Hub |
| 72 | o1 | OpenAI | 75.6% | 2026-07-20 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 73 | Grok-3 mini | xAI | 75.4% | 2025-05-26 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 74 | Claude Sonnet 4 | Anthropic | 74.6% | 2025-05-26 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 75 | Claude 3.7 Sonnet | Anthropic | 74.5% | 2025-05-26 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 76 | o3-mini | OpenAI | 73.2% | 2026-07-20 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 77 | GPT-5 mini | OpenAI | 72.8% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 78 | Claude Opus 4 | Anthropic | 72.7% | 2025-05-22 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 79 | Qwen3-Max | Alibaba | 72.6% | 2025-10-06 | independent | Epoch AI Benchmarking Hub |
| 80 | Qwen3 235B-A22Bopen | Alibaba | 70.7% | 2025-06-03 | independent | Epoch AI Benchmarking Hub |
| 81 | DeepSeek-R1open | DeepSeek | 69.2% | 2025-05-26 | independent | Epoch AI Benchmarking Hub |
| 82 | GPT-4.5 | OpenAI | 68.7% | 2025-02-28 | independent | Epoch AI Benchmarking Hub |
| 83 | GPT-5.4 Nano | OpenAI | 68.7% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 84 | DeepSeek-V3 (Mar 2025)open | DeepSeek | 67.6% | 2025-04-01 | independent | Epoch AI Benchmarking Hub |
| 85 | Grok 3 | xAI | 67.6% | 2025-05-26 | independent | Epoch AI Benchmarking Hub |
| 86 | Llama 4 Maverickopen | Meta AI | 67.0% | 2025-04-08 | independent | Epoch AI Benchmarking Hub |
| 87 | GPT-4.1 | OpenAI | 66.9% | 2025-04-14 | independent | Epoch AI Benchmarking Hub |
| 88 | Gemini 2.5 Pro (May 2025) | Google DeepMind | 66.7% | 2025-06-03 | independent | Epoch AI Benchmarking Hub |
| 89 | Claude Haiku 4.5 | Anthropic | 65.8% | 2025-10-22 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 90 | GPT-4.1 mini | OpenAI | 65.8% | 2025-04-14 | independent | Epoch AI Benchmarking Hub |
| 91 | Gemini 2.0 Pro | Google DeepMind | 65.7% | 2025-02-07 | independent | Epoch AI Benchmarking Hub |
| 92 | QWQ-Plus | Alibaba | 65.4% | 2025-04-11 | independent | Epoch AI Benchmarking Hub |
| 93 | Gemini 2.0 Flash | Google DeepMind | 64.1% | 2025-02-06 | independent | Epoch AI Benchmarking Hub |
| 94 | o1-mini | OpenAI | 60.9% | 2025-02-13 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 95 | GPT-5 nano | OpenAI | 60.7% | 2026-07-13 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 96 | Mistral Medium 3 | Mistral | 59.5% | 2025-05-07 | independent | Epoch AI Benchmarking Hub |
| 97 | Gemini 2.0 Flash Thinking | Google DeepMind | 57.1% | 2025-02-06 | independent | Epoch AI Benchmarking Hub |
| 98 | DeepSeek-V3open | DeepSeek | 56.5% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 99 | Magistral Small 1.0open | Mistral | 56.1% | 2026-08-06 | independent | Epoch AI Benchmarking Hub |
| 100 | Phi-4open | Microsoft | 56.1% | 2025-01-31 | independent | Epoch AI Benchmarking Hub |
| 101 | Qwen2.5-Max | Alibaba | 56.1% | 2025-04-01 | independent | Epoch AI Benchmarking Hub |
| 102 | DeepSeek-R1-Distill-Llama-70Bopen | DeepSeek | 55.7% | 2025-03-10 | independent | Epoch AI Benchmarking Hub |
| 103 | Claude 3.5 Sonnet (October 2024) | Anthropic | 55.3% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 104 | Claude 3.5 Sonnet | Anthropic | 54.0% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 105 | Grok-2 | xAI | 53.8% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 106 | Llama 4 Scoutopen | Meta AI | 51.8% | 2025-04-08 | independent | Epoch AI Benchmarking Hub |
| 107 | Gemini 1.5 Pro | Google DeepMind | 51.5% | 2025-01-27 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 108 | Llama 3.1-405Bopen | Meta AI | 50.9% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 109 | o1-preview | OpenAI | 50.3% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 110 | Mistral Large 2open | Mistral | 50.2% | 2025-02-25 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 111 | Qwen2.5-72Bopen | Alibaba | 49.1% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 112 | GPT-4.1 nano | OpenAI | 48.9% | 2025-04-14 | independent | Epoch AI Benchmarking Hub |
| 113 | GPT-4o | OpenAI | 48.7% | 2025-02-05 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 114 | Qwen Plus | Alibaba | 48.1% | 2025-04-07 | independent | Epoch AI Benchmarking Hub |
| 115 | Mistral Small 3.1open | Mistral | 47.5% | 2025-03-18 | independent | Epoch AI Benchmarking Hub |
| 116 | Llama 3.3 70Bopen | Meta AI | 47.4% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 117 | Claude 3 Opus | Anthropic | 47.2% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 118 | GPT-4 Turbo (Apr 2024) | OpenAI | 46.6% | 2025-02-27 | independent | Epoch AI Benchmarking Hub |
| 119 | Tulu 3 (Tülu 3) 70Bopen | Allen Institute | 46.3% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 120 | Qwen2.5-32Bopen | Alibaba | 46.1% | 2025-01-30 | independent | Epoch AI Benchmarking Hub |
| 121 | gpt-oss-20bopen | OpenAI | 46.0% | 2026-08-06 | independent | Epoch AI Benchmarking Hub |
| 122 | Mistral Small 3open | Mistral | 45.3% | 2025-01-30 | independent | Epoch AI Benchmarking Hub |
| 123 | DeepSeek-R1-Distill-Qwen-14Bopen | DeepSeek | 44.7% | 2025-03-10 | independent | Epoch AI Benchmarking Hub |
| 124 | Llama 3.1-70Bopen | Meta AI | 44.2% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 125 | Gemma 3 27Bopen | Google DeepMind | 43.9% | 2026-08-06 | independent | Epoch AI Benchmarking Hub |
| 126 | Gemini 1.5 Flash | Google DeepMind | 43.8% | 2025-01-27 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 127 | WizardLM-2 8x22Bopen | Microsoft | 43.4% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 128 | GPT-4 Turbo (Nov 2023) | OpenAI | 42.3% | 2025-01-27 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 129 | Qwen-Turbo | Alibaba | 41.8% | 2025-04-07 | independent | Epoch AI Benchmarking Hub |
| 130 | Llama 3.2 90Bopen | Meta AI | 41.0% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 131 | Qwen2-72Bopen | Alibaba | 40.8% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 132 | Claude 3 Sonnet | Anthropic | 40.6% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 133 | Llama 3-70Bopen | Meta AI | 40.6% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 134 | Mistral Large | Mistral | 38.8% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 135 | Claude 3.5 Haiku | Anthropic | 38.1% | 2025-03-12 | independent | Epoch AI Benchmarking Hub |
| 136 | GPT-4o mini | OpenAI | 37.7% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 137 | Hermes 2 Theta Llama-3 70Bopen | Other | 37.5% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 138 | Gemma 2 27Bopen | Google DeepMind | 36.5% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 139 | Claude 3 Haiku | Anthropic | 36.3% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 140 | GPT-4 (Mar 2023) | OpenAI | 35.7% | 2025-10-23 | independent | Epoch AI Benchmarking Hub |
| 141 | Claude 2 | Anthropic | 34.7% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 142 | Mixtral 8x22Bopen | Mistral | 34.1% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 143 | Gemini 1.0 Pro | Google DeepMind | 34.0% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 144 | Eurus-2-7B-PRIMEopen | Other | 33.9% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 145 | Claude 2.1 | Anthropic | 33.0% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 146 | Gemini 1.5 Flash 8B | Google DeepMind | 33.0% | 2025-03-12 | independent | Epoch AI Benchmarking Hub |
| 147 | DBRXopen | Other | 32.9% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 148 | Yi-1.5-34Bopen | Other | 32.0% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 149 | GPT-4 (Jun 2023) | OpenAI | 30.7% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 150 | Qwen1.5-32Bopen | Alibaba | 30.7% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 151 | Mixtral 8x7Bopen | Mistral | 30.2% | 2025-01-27 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 152 | Mistral NeMoopen | Mistral | 29.9% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 153 | Qwen1.5-72Bopen | Alibaba | 28.8% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 154 | GPT-3.5 Turbo | OpenAI | 27.6% | 2025-01-27 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 155 | phi-3-medium 14Bopen | Microsoft | 27.6% | 2025-01-31 | independent | Epoch AI Benchmarking Hub |
| 156 | Gemma 2 9Bopen | Google DeepMind | 27.5% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 157 | Ministral 8Bopen | Mistral | 27.1% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 158 | Llama 2-70Bopen | Meta AI | 26.3% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 159 | Llama 3-8Bopen | Meta AI | 26.1% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 160 | Llama 3.1-8Bopen | Meta AI | 25.9% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 161 | Ministral 3B | Mistral | 25.3% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 162 | DeepSeek LLM 67Bopen | DeepSeek | 24.6% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 163 | Yi-34Bopen | Other | 14.7% | 2025-01-27 | independent | Epoch AI Benchmarking Hub |
| 164 | Mistral 7Bopen | Mistral | 14.2% | 2025-01-27 | independent · 2 runs | Epoch AI Benchmarking Hub |