Short factual questions with a single unambiguous answer, cleaned of mislabelled items from the original SimpleQA set.
Factual recall and, above all, the willingness to say nothing rather than hallucinate.
Models with web access have a large advantage. Scores here are without tools unless noted.
| # | Model | Lab | Score | Measured | Method | Source |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro | Google DeepMind | 77.3% | 2026-02-19 | independent | Epoch AI Benchmarking Hub |
| 2 | Gemini 3 Pro | Google DeepMind | 72.9% | 2025-12-08 | independent | Epoch AI Benchmarking Hub |
| 3 | GPT-5.6 Sol | OpenAI | 71.6% | 2026-07-09 | independent | Epoch AI Benchmarking Hub |
| 4 | Gemini 3.7 Flash | Google DeepMind | 71.2% | 2026-08-14 | independent | Epoch AI Benchmarking Hub |
| 5 | Gemini 3.6 Flash | Google DeepMind | 68.7% | 2026-08-02 | independent | Epoch AI Benchmarking Hub |
| 6 | Gemini 3.5 Flash | Google DeepMind | 68.4% | 2026-05-27 | independent | Epoch AI Benchmarking Hub |
| 7 | Claude Fable 5 | Anthropic | 68.3% | 2026-06-09 | independent | Epoch AI Benchmarking Hub |
| 8 | Qwen3-Max | Alibaba | 67.5% | 2025-12-10 | independent | Epoch AI Benchmarking Hub |
| 9 | Gemini 3 Flash | Google DeepMind | 67.4% | 2025-12-17 | independent | Epoch AI Benchmarking Hub |
| 10 | Muse Spark | Meta AI | 66.3% | 2026-04-08 | independent | Epoch AI Benchmarking Hub |
| 11 | GPT-5.5 Pro | OpenAI | 64.5% | 2026-04-24 | independent | Epoch AI Benchmarking Hub |
| 12 | GPT-5.5 | OpenAI | 63.1% | 2026-04-24 | independent | Epoch AI Benchmarking Hub |
| 13 | Qwen3.7-Max | Alibaba | 58.5% | 2026-06-13 | independent | Epoch AI Benchmarking Hub |
| 14 | DeepSeek-V4-Proopen | DeepSeek | 57.0% | 2026-06-16 | independent | Epoch AI Benchmarking Hub |
| 15 | Qwen 3.6 Max (Preview) | Alibaba | 56.9% | 2026-05-15 | independent | Epoch AI Benchmarking Hub |
| 16 | Claude Opus 5 | Anthropic | 56.7% | 2026-07-24 | independent | Epoch AI Benchmarking Hub |
| 17 | Gemini 2.5 Pro (Jun 2025) | Google DeepMind | 56.0% | 2025-12-08 | independent | Epoch AI Benchmarking Hub |
| 18 | Grok 4.6 | xAI | 53.7% | 2026-08-14 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 19 | Grok 4.5 | xAI | 53.5% | 2026-07-08 | independent | Epoch AI Benchmarking Hub |
| 20 | o3 | OpenAI | 53.0% | 2025-12-09 | independent | Epoch AI Benchmarking Hub |
| 21 | Claude Opus 4.7 | Anthropic | 50.6% | 2026-04-16 | independent | Epoch AI Benchmarking Hub |
| 22 | GPT-5 | OpenAI | 50.6% | 2025-12-09 | independent | Epoch AI Benchmarking Hub |
| 23 | Qwen3-235B-A22B (Jul 2025)open | Alibaba | 50.1% | 2025-12-11 | independent | Epoch AI Benchmarking Hub |
| 24 | Qwen 3.6 Plus | Alibaba | 49.1% | 2026-05-12 | independent | Epoch AI Benchmarking Hub |
| 25 | GPT-5.1 | OpenAI | 48.9% | 2025-12-09 | independent | Epoch AI Benchmarking Hub |
| 26 | Grok 4 | xAI | 47.9% | 2025-12-08 | independent | Epoch AI Benchmarking Hub |
| 27 | GPT-5.4 Pro | OpenAI | 47.8% | 2026-03-19 | independent | Epoch AI Benchmarking Hub |
| 28 | Qwen 3.8 Max | Alibaba | 46.3% | 2026-08-04 | independent | Epoch AI Benchmarking Hub |
| 29 | GPT-5.4 | OpenAI | 44.8% | 2026-03-06 | independent | Epoch AI Benchmarking Hub |
| 30 | Claude Opus 4.6 | Anthropic | 43.5% | 2026-02-13 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 31 | GPT-5.6 Terra | OpenAI | 43.1% | 2026-07-09 | independent | Epoch AI Benchmarking Hub |
| 32 | Kimi K3open | Moonshot AI | 42.7% | 2026-07-16 | independent | Epoch AI Benchmarking Hub |
| 33 | Claude Opus 4.5 | Anthropic | 41.8% | 2025-12-09 | independent | Epoch AI Benchmarking Hub |
| 34 | GPT-5.6 Luna | OpenAI | 41.7% | 2026-07-09 | independent | Epoch AI Benchmarking Hub |
| 35 | Inklingopen | Thinking Machines | 40.2% | 2026-08-05 | independent | Epoch AI Benchmarking Hub |
| 36 | Claude Opus 4.8 | Anthropic | 39.5% | 2026-05-29 | independent | Epoch AI Benchmarking Hub |
| 37 | Kimi K2.7 Codeopen | Moonshot AI | 39.2% | 2026-06-12 | independent | Epoch AI Benchmarking Hub |
| 38 | Kimi K2.6open | Moonshot AI | 38.7% | 2026-05-01 | independent | Epoch AI Benchmarking Hub |
| 39 | GLM-5.2open | Z.ai | 38.1% | 2026-06-19 | independent | Epoch AI Benchmarking Hub |
| 40 | GPT-5.5 Instant | OpenAI | 38.0% | 2026-08-02 | independent | Epoch AI Benchmarking Hub |
| 41 | Grok 4.3 Beta | xAI | 38.0% | 2026-06-16 | independent | Epoch AI Benchmarking Hub |
| 42 | GLM-5.1open | Z.ai | 37.3% | 2026-05-01 | independent | Epoch AI Benchmarking Hub |
| 43 | GPT-5.2 | OpenAI | 36.8% | 2025-12-11 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 44 | Grok 4.20 | xAI | 35.1% | 2026-07-13 | independent | Epoch AI Benchmarking Hub |
| 45 | Claude Opus 4.1 | Anthropic | 34.8% | 2025-12-09 | independent | Epoch AI Benchmarking Hub |
| 46 | DeepSeek V4 Flash 0731open | DeepSeek | 34.7% | 2026-08-02 | independent | Epoch AI Benchmarking Hub |
| 47 | Kimi K2.5open | Moonshot AI | 33.9% | 2026-01-28 | independent | Epoch AI Benchmarking Hub |
| 48 | Kimi K2 Thinkingopen | Moonshot AI | 31.6% | 2025-12-10 | independent | Epoch AI Benchmarking Hub |
| 49 | GLM-4.7open | Z.ai | 31.5% | 2026-01-30 | independent | Epoch AI Benchmarking Hub |
| 50 | Claude Sonnet 4.6 | Anthropic | 29.0% | 2026-02-21 | independent | Epoch AI Benchmarking Hub |
| 51 | GPT-5.4 Mini | OpenAI | 28.6% | 2026-04-15 | independent | Epoch AI Benchmarking Hub |
| 52 | DeepSeek-V3.2open | DeepSeek | 27.5% | 2025-12-16 | independent | Epoch AI Benchmarking Hub |
| 53 | DeepSeek-R1 (May 2025)open | DeepSeek | 27.4% | 2025-12-08 | independent | Epoch AI Benchmarking Hub |
| 54 | Qwen 3.5 Plus (hosted 397B-A17B) | Alibaba | 26.0% | 2026-05-13 | independent | Epoch AI Benchmarking Hub |
| 55 | Claude Sonnet 5 | Anthropic | 23.7% | 2026-07-01 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 56 | o4-mini | OpenAI | 22.9% | 2026-07-11 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 57 | Qwen 3.6 Flash | Alibaba | 21.2% | 2026-05-12 | independent | Epoch AI Benchmarking Hub |
| 58 | Grok-3 mini | xAI | 21.1% | 2025-12-09 | independent | Epoch AI Benchmarking Hub |
| 59 | GPT-5 mini | OpenAI | 21.0% | 2025-12-09 | independent | Epoch AI Benchmarking Hub |
| 60 | Qwen 3.5 Flash (hosted 35B-A3B) | Alibaba | 19.8% | 2026-05-11 | independent | Epoch AI Benchmarking Hub |
| 61 | Inkling-Smallopen | Thinking Machines | 19.5% | 2026-08-14 | independent | Epoch AI Benchmarking Hub |
| 62 | Claude Sonnet 4.5 | Anthropic | 18.3% | 2025-12-09 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 63 | gpt-oss-120bopen | OpenAI | 13.9% | 2025-12-15 | independent | Epoch AI Benchmarking Hub |
| 64 | GPT-5 nano | OpenAI | 12.2% | 2025-12-09 | independent | Epoch AI Benchmarking Hub |
| 65 | GPT-5.4 Nano | OpenAI | 12.0% | 2026-04-14 | independent | Epoch AI Benchmarking Hub |
| 66 | Gemma 4 31B ITopen | Google DeepMind | 9.6% | 2026-04-04 | independent | Epoch AI Benchmarking Hub |
| 67 | Claude 3.5 Haiku | Anthropic | 6.7% | 2025-11-28 | independent | Epoch AI Benchmarking Hub |
| 68 | Claude Haiku 4.5 | Anthropic | 5.9% | 2025-12-09 | independent | Epoch AI Benchmarking Hub |