← All benchmarks

SimpleQA Verified

Knowledge unit: % independent 68 models scored counts toward composite (weight 0.75)

Short factual questions with a single unambiguous answer, cleaned of mislabelled items from the original SimpleQA set.

What it measures

Factual recall and, above all, the willingness to say nothing rather than hallucinate.

How to read it

Models with web access have a large advantage. Scores here are without tools unless noted.

Source

Epoch AI Benchmarking Hub

Full ranking

#ModelLabScoreMeasuredMethodSource
1 Gemini 3.1 Pro Google DeepMind 77.3% 2026-02-19 independent Epoch AI Benchmarking Hub
2 Gemini 3 Pro Google DeepMind 72.9% 2025-12-08 independent Epoch AI Benchmarking Hub
3 GPT-5.6 Sol OpenAI 71.6% 2026-07-09 independent Epoch AI Benchmarking Hub
4 Gemini 3.7 Flash Google DeepMind 71.2% 2026-08-14 independent Epoch AI Benchmarking Hub
5 Gemini 3.6 Flash Google DeepMind 68.7% 2026-08-02 independent Epoch AI Benchmarking Hub
6 Gemini 3.5 Flash Google DeepMind 68.4% 2026-05-27 independent Epoch AI Benchmarking Hub
7 Claude Fable 5 Anthropic 68.3% 2026-06-09 independent Epoch AI Benchmarking Hub
8 Qwen3-Max Alibaba 67.5% 2025-12-10 independent Epoch AI Benchmarking Hub
9 Gemini 3 Flash Google DeepMind 67.4% 2025-12-17 independent Epoch AI Benchmarking Hub
10 Muse Spark Meta AI 66.3% 2026-04-08 independent Epoch AI Benchmarking Hub
11 GPT-5.5 Pro OpenAI 64.5% 2026-04-24 independent Epoch AI Benchmarking Hub
12 GPT-5.5 OpenAI 63.1% 2026-04-24 independent Epoch AI Benchmarking Hub
13 Qwen3.7-Max Alibaba 58.5% 2026-06-13 independent Epoch AI Benchmarking Hub
14 DeepSeek-V4-Proopen DeepSeek 57.0% 2026-06-16 independent Epoch AI Benchmarking Hub
15 Qwen 3.6 Max (Preview) Alibaba 56.9% 2026-05-15 independent Epoch AI Benchmarking Hub
16 Claude Opus 5 Anthropic 56.7% 2026-07-24 independent Epoch AI Benchmarking Hub
17 Gemini 2.5 Pro (Jun 2025) Google DeepMind 56.0% 2025-12-08 independent Epoch AI Benchmarking Hub
18 Grok 4.6 xAI 53.7% 2026-08-14 independent · 2 runs Epoch AI Benchmarking Hub
19 Grok 4.5 xAI 53.5% 2026-07-08 independent Epoch AI Benchmarking Hub
20 o3 OpenAI 53.0% 2025-12-09 independent Epoch AI Benchmarking Hub
21 Claude Opus 4.7 Anthropic 50.6% 2026-04-16 independent Epoch AI Benchmarking Hub
22 GPT-5 OpenAI 50.6% 2025-12-09 independent Epoch AI Benchmarking Hub
23 Qwen3-235B-A22B (Jul 2025)open Alibaba 50.1% 2025-12-11 independent Epoch AI Benchmarking Hub
24 Qwen 3.6 Plus Alibaba 49.1% 2026-05-12 independent Epoch AI Benchmarking Hub
25 GPT-5.1 OpenAI 48.9% 2025-12-09 independent Epoch AI Benchmarking Hub
26 Grok 4 xAI 47.9% 2025-12-08 independent Epoch AI Benchmarking Hub
27 GPT-5.4 Pro OpenAI 47.8% 2026-03-19 independent Epoch AI Benchmarking Hub
28 Qwen 3.8 Max Alibaba 46.3% 2026-08-04 independent Epoch AI Benchmarking Hub
29 GPT-5.4 OpenAI 44.8% 2026-03-06 independent Epoch AI Benchmarking Hub
30 Claude Opus 4.6 Anthropic 43.5% 2026-02-13 independent · 3 runs Epoch AI Benchmarking Hub
31 GPT-5.6 Terra OpenAI 43.1% 2026-07-09 independent Epoch AI Benchmarking Hub
32 Kimi K3open Moonshot AI 42.7% 2026-07-16 independent Epoch AI Benchmarking Hub
33 Claude Opus 4.5 Anthropic 41.8% 2025-12-09 independent Epoch AI Benchmarking Hub
34 GPT-5.6 Luna OpenAI 41.7% 2026-07-09 independent Epoch AI Benchmarking Hub
35 Inklingopen Thinking Machines 40.2% 2026-08-05 independent Epoch AI Benchmarking Hub
36 Claude Opus 4.8 Anthropic 39.5% 2026-05-29 independent Epoch AI Benchmarking Hub
37 Kimi K2.7 Codeopen Moonshot AI 39.2% 2026-06-12 independent Epoch AI Benchmarking Hub
38 Kimi K2.6open Moonshot AI 38.7% 2026-05-01 independent Epoch AI Benchmarking Hub
39 GLM-5.2open Z.ai 38.1% 2026-06-19 independent Epoch AI Benchmarking Hub
40 GPT-5.5 Instant OpenAI 38.0% 2026-08-02 independent Epoch AI Benchmarking Hub
41 Grok 4.3 Beta xAI 38.0% 2026-06-16 independent Epoch AI Benchmarking Hub
42 GLM-5.1open Z.ai 37.3% 2026-05-01 independent Epoch AI Benchmarking Hub
43 GPT-5.2 OpenAI 36.8% 2025-12-11 independent · 4 runs Epoch AI Benchmarking Hub
44 Grok 4.20 xAI 35.1% 2026-07-13 independent Epoch AI Benchmarking Hub
45 Claude Opus 4.1 Anthropic 34.8% 2025-12-09 independent Epoch AI Benchmarking Hub
46 DeepSeek V4 Flash 0731open DeepSeek 34.7% 2026-08-02 independent Epoch AI Benchmarking Hub
47 Kimi K2.5open Moonshot AI 33.9% 2026-01-28 independent Epoch AI Benchmarking Hub
48 Kimi K2 Thinkingopen Moonshot AI 31.6% 2025-12-10 independent Epoch AI Benchmarking Hub
49 GLM-4.7open Z.ai 31.5% 2026-01-30 independent Epoch AI Benchmarking Hub
50 Claude Sonnet 4.6 Anthropic 29.0% 2026-02-21 independent Epoch AI Benchmarking Hub
51 GPT-5.4 Mini OpenAI 28.6% 2026-04-15 independent Epoch AI Benchmarking Hub
52 DeepSeek-V3.2open DeepSeek 27.5% 2025-12-16 independent Epoch AI Benchmarking Hub
53 DeepSeek-R1 (May 2025)open DeepSeek 27.4% 2025-12-08 independent Epoch AI Benchmarking Hub
54 Qwen 3.5 Plus (hosted 397B-A17B) Alibaba 26.0% 2026-05-13 independent Epoch AI Benchmarking Hub
55 Claude Sonnet 5 Anthropic 23.7% 2026-07-01 independent · 2 runs Epoch AI Benchmarking Hub
56 o4-mini OpenAI 22.9% 2026-07-11 independent · 2 runs Epoch AI Benchmarking Hub
57 Qwen 3.6 Flash Alibaba 21.2% 2026-05-12 independent Epoch AI Benchmarking Hub
58 Grok-3 mini xAI 21.1% 2025-12-09 independent Epoch AI Benchmarking Hub
59 GPT-5 mini OpenAI 21.0% 2025-12-09 independent Epoch AI Benchmarking Hub
60 Qwen 3.5 Flash (hosted 35B-A3B) Alibaba 19.8% 2026-05-11 independent Epoch AI Benchmarking Hub
61 Inkling-Smallopen Thinking Machines 19.5% 2026-08-14 independent Epoch AI Benchmarking Hub
62 Claude Sonnet 4.5 Anthropic 18.3% 2025-12-09 independent · 2 runs Epoch AI Benchmarking Hub
63 gpt-oss-120bopen OpenAI 13.9% 2025-12-15 independent Epoch AI Benchmarking Hub
64 GPT-5 nano OpenAI 12.2% 2025-12-09 independent Epoch AI Benchmarking Hub
65 GPT-5.4 Nano OpenAI 12.0% 2026-04-14 independent Epoch AI Benchmarking Hub
66 Gemma 4 31B ITopen Google DeepMind 9.6% 2026-04-04 independent Epoch AI Benchmarking Hub
67 Claude 3.5 Haiku Anthropic 6.7% 2025-11-28 independent Epoch AI Benchmarking Hub
68 Claude Haiku 4.5 Anthropic 5.9% 2025-12-09 independent Epoch AI Benchmarking Hub