← All benchmarks

Mystery Game Puzzles

Games unit: % independent 35 models scored not in composite

Deduction puzzles set in unfamiliar rule systems the model has to infer from interaction rather than recall.

What it measures

Learning rules on the fly and reasoning under incomplete information.

How to read it

Very hard for current models. Most scores are close to the floor.

Source

Epoch AI Benchmarking Hub

Full ranking

#ModelLabScoreMeasuredMethodSource
1 GPT-5.6 Sol OpenAI 58.0% 2026-07-28 independent Epoch AI Benchmarking Hub
2 Claude Fable 5 Anthropic 52.0% 2026-07-31 independent Epoch AI Benchmarking Hub
3 Claude Opus 5 Anthropic 48.0% 2026-08-06 independent · 2 runs Epoch AI Benchmarking Hub
4 GPT-5.5 OpenAI 46.7% 2026-08-06 independent · 3 runs Epoch AI Benchmarking Hub
5 Qwen 3.8 Max Alibaba 38.0% 2026-08-05 independent Epoch AI Benchmarking Hub
6 Gemini 3.7 Flash Google DeepMind 37.0% 2026-08-14 independent Epoch AI Benchmarking Hub
7 GPT-5.4 OpenAI 37.0% 2026-07-24 independent Epoch AI Benchmarking Hub
8 GPT-5.6 Terra OpenAI 35.0% 2026-07-28 independent Epoch AI Benchmarking Hub
9 DeepSeek V4 Flash 0731open DeepSeek 34.0% 2026-08-05 independent Epoch AI Benchmarking Hub
10 Grok 4.6 xAI 34.0% 2026-08-14 independent Epoch AI Benchmarking Hub
11 Claude Opus 4.8 Anthropic 33.5% 2026-07-26 independent · 2 runs Epoch AI Benchmarking Hub
12 Qwen3.7-Max Alibaba 32.0% 2026-07-28 independent Epoch AI Benchmarking Hub
13 Gemini 3.1 Pro Google DeepMind 31.7% 2026-08-06 independent · 3 runs Epoch AI Benchmarking Hub
14 Gemini 3.5 Flash Google DeepMind 30.0% 2026-08-05 independent · 2 runs Epoch AI Benchmarking Hub
15 Qwen 3.6 Plus Alibaba 27.0% 2026-08-05 independent Epoch AI Benchmarking Hub
16 Gemini 3.6 Flash Google DeepMind 26.0% 2026-08-05 independent · 3 runs Epoch AI Benchmarking Hub
17 Kimi K3open Moonshot AI 26.0% 2026-07-29 independent Epoch AI Benchmarking Hub
18 Claude Sonnet 5 Anthropic 25.5% 2026-08-06 independent · 2 runs Epoch AI Benchmarking Hub
19 Gemini 3 Flash Google DeepMind 23.7% 2026-08-05 independent · 3 runs Epoch AI Benchmarking Hub
20 GPT-5 OpenAI 23.0% 2026-08-05 independent Epoch AI Benchmarking Hub
21 GPT-5.2 OpenAI 23.0% 2026-08-06 independent Epoch AI Benchmarking Hub
22 Claude Opus 4.5 Anthropic 22.0% 2026-07-25 independent Epoch AI Benchmarking Hub
23 Claude Opus 4.1 Anthropic 21.0% 2026-07-25 independent Epoch AI Benchmarking Hub
24 GPT-5.6 Luna OpenAI 21.0% 2026-07-28 independent Epoch AI Benchmarking Hub
25 Claude Opus 4.7 Anthropic 20.5% 2026-08-06 independent · 2 runs Epoch AI Benchmarking Hub
26 Kimi K2.6open Moonshot AI 18.0% 2026-07-17 independent Epoch AI Benchmarking Hub
27 Claude Sonnet 4.5 Anthropic 17.0% 2026-07-25 independent Epoch AI Benchmarking Hub
28 Qwen 3.6 Max (Preview) Alibaba 16.0% 2026-08-05 independent Epoch AI Benchmarking Hub
29 Claude Opus 4.6 Anthropic 15.7% 2026-08-06 independent · 3 runs Epoch AI Benchmarking Hub
30 Gemini 3.5 Flash-Lite Google DeepMind 15.5% 2026-08-05 independent · 2 runs Epoch AI Benchmarking Hub
31 Qwen 3.5 Plus (hosted 397B-A17B) Alibaba 15.5% 2026-08-05 independent · 2 runs Epoch AI Benchmarking Hub
32 Claude Sonnet 4.6 Anthropic 15.0% 2026-08-06 independent · 2 runs Epoch AI Benchmarking Hub
33 Qwen 3.5 Flash (hosted 35B-A3B) Alibaba 15.0% 2026-08-05 independent Epoch AI Benchmarking Hub
34 Qwen 3.6 Flash Alibaba 14.0% 2026-08-05 independent Epoch AI Benchmarking Hub
35 Inkling-Smallopen Thinking Machines 6.0% 2026-08-15 independent Epoch AI Benchmarking Hub