Deduction puzzles set in unfamiliar rule systems the model has to infer from interaction rather than recall.
Learning rules on the fly and reasoning under incomplete information.
Very hard for current models. Most scores are close to the floor.
| # | Model | Lab | Score | Measured | Method | Source |
|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Sol | OpenAI | 58.0% | 2026-07-28 | independent | Epoch AI Benchmarking Hub |
| 2 | Claude Fable 5 | Anthropic | 52.0% | 2026-07-31 | independent | Epoch AI Benchmarking Hub |
| 3 | Claude Opus 5 | Anthropic | 48.0% | 2026-08-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 4 | GPT-5.5 | OpenAI | 46.7% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 5 | Qwen 3.8 Max | Alibaba | 38.0% | 2026-08-05 | independent | Epoch AI Benchmarking Hub |
| 6 | Gemini 3.7 Flash | Google DeepMind | 37.0% | 2026-08-14 | independent | Epoch AI Benchmarking Hub |
| 7 | GPT-5.4 | OpenAI | 37.0% | 2026-07-24 | independent | Epoch AI Benchmarking Hub |
| 8 | GPT-5.6 Terra | OpenAI | 35.0% | 2026-07-28 | independent | Epoch AI Benchmarking Hub |
| 9 | DeepSeek V4 Flash 0731open | DeepSeek | 34.0% | 2026-08-05 | independent | Epoch AI Benchmarking Hub |
| 10 | Grok 4.6 | xAI | 34.0% | 2026-08-14 | independent | Epoch AI Benchmarking Hub |
| 11 | Claude Opus 4.8 | Anthropic | 33.5% | 2026-07-26 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 12 | Qwen3.7-Max | Alibaba | 32.0% | 2026-07-28 | independent | Epoch AI Benchmarking Hub |
| 13 | Gemini 3.1 Pro | Google DeepMind | 31.7% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 14 | Gemini 3.5 Flash | Google DeepMind | 30.0% | 2026-08-05 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 15 | Qwen 3.6 Plus | Alibaba | 27.0% | 2026-08-05 | independent | Epoch AI Benchmarking Hub |
| 16 | Gemini 3.6 Flash | Google DeepMind | 26.0% | 2026-08-05 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 17 | Kimi K3open | Moonshot AI | 26.0% | 2026-07-29 | independent | Epoch AI Benchmarking Hub |
| 18 | Claude Sonnet 5 | Anthropic | 25.5% | 2026-08-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 19 | Gemini 3 Flash | Google DeepMind | 23.7% | 2026-08-05 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 20 | GPT-5 | OpenAI | 23.0% | 2026-08-05 | independent | Epoch AI Benchmarking Hub |
| 21 | GPT-5.2 | OpenAI | 23.0% | 2026-08-06 | independent | Epoch AI Benchmarking Hub |
| 22 | Claude Opus 4.5 | Anthropic | 22.0% | 2026-07-25 | independent | Epoch AI Benchmarking Hub |
| 23 | Claude Opus 4.1 | Anthropic | 21.0% | 2026-07-25 | independent | Epoch AI Benchmarking Hub |
| 24 | GPT-5.6 Luna | OpenAI | 21.0% | 2026-07-28 | independent | Epoch AI Benchmarking Hub |
| 25 | Claude Opus 4.7 | Anthropic | 20.5% | 2026-08-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 26 | Kimi K2.6open | Moonshot AI | 18.0% | 2026-07-17 | independent | Epoch AI Benchmarking Hub |
| 27 | Claude Sonnet 4.5 | Anthropic | 17.0% | 2026-07-25 | independent | Epoch AI Benchmarking Hub |
| 28 | Qwen 3.6 Max (Preview) | Alibaba | 16.0% | 2026-08-05 | independent | Epoch AI Benchmarking Hub |
| 29 | Claude Opus 4.6 | Anthropic | 15.7% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 30 | Gemini 3.5 Flash-Lite | Google DeepMind | 15.5% | 2026-08-05 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 31 | Qwen 3.5 Plus (hosted 397B-A17B) | Alibaba | 15.5% | 2026-08-05 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 32 | Claude Sonnet 4.6 | Anthropic | 15.0% | 2026-08-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 33 | Qwen 3.5 Flash (hosted 35B-A3B) | Alibaba | 15.0% | 2026-08-05 | independent | Epoch AI Benchmarking Hub |
| 34 | Qwen 3.6 Flash | Alibaba | 14.0% | 2026-08-05 | independent | Epoch AI Benchmarking Hub |
| 35 | Inkling-Smallopen | Thinking Machines | 6.0% | 2026-08-15 | independent | Epoch AI Benchmarking Hub |