Code tasks built to mirror the structure of well-known problems while changing the underlying semantics, so memorised solutions fail.
Whether coding ability generalises beyond memorised patterns.
Few models evaluated. Treat as directional.
| # | Model | Lab | Score | Measured | Method | Source |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 63.9% | 2026-08-10 | independent | Epoch AI Benchmarking Hub |
| 2 | Claude Opus 4.7 | Anthropic | 31.1% | 2026-08-12 | independent | Epoch AI Benchmarking Hub |
| 3 | GPT-5.6 Sol | OpenAI | 20.0% | 2026-08-10 | independent | Epoch AI Benchmarking Hub |
| 4 | GPT-5.4 | OpenAI | 15.6% | 2026-08-10 | independent | Epoch AI Benchmarking Hub |
| 5 | GPT-5.5 | OpenAI | 10.0% | 2026-08-10 | independent | Epoch AI Benchmarking Hub |
| 6 | Gemini 3.1 Pro | Google DeepMind | 8.9% | 2026-08-10 | independent | Epoch AI Benchmarking Hub |