Mock American Invitational Mathematics Examination problems written for the OTIS olympiad program, used as a contest-math benchmark that is not part of the public AIME sets.
Competition mathematics at olympiad-qualifier level.
Close to saturation for frontier reasoning models, so it separates the top of the field less and less.
| # | Model | Lab | Score | Measured | Method | Source |
|---|---|---|---|---|---|---|
| 1 | GPT-5.5 Pro | OpenAI | 100.0% | 2026-04-24 | independent | Epoch AI Benchmarking Hub |
| 2 | Qwen 3.8 Max | Alibaba | 99.4% | 2026-08-04 | independent | Epoch AI Benchmarking Hub |
| 3 | Claude Fable 5 | Anthropic | 99.2% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 4 | Grok 4.6 | xAI | 98.5% | 2026-08-14 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 5 | Grok 4.5 | xAI | 97.8% | 2026-07-08 | independent | Epoch AI Benchmarking Hub |
| 6 | Gemini 3.7 Flash | Google DeepMind | 97.2% | 2026-08-14 | independent | Epoch AI Benchmarking Hub |
| 7 | Claude Opus 5 | Anthropic | 96.7% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 8 | Kimi K2.6open | Moonshot AI | 96.1% | 2026-05-02 | independent | Epoch AI Benchmarking Hub |
| 9 | Gemini 3.1 Pro | Google DeepMind | 95.6% | 2026-08-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 10 | Kimi K2.7 Codeopen | Moonshot AI | 95.6% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 11 | Qwen3.7-Max | Alibaba | 95.6% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 12 | DeepSeek V4 Flash 0731open | DeepSeek | 94.4% | 2026-08-02 | independent | Epoch AI Benchmarking Hub |
| 13 | Gemini 3 Flash | Google DeepMind | 94.2% | 2026-08-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 14 | Claude Opus 4.8 | Anthropic | 93.5% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 15 | GLM-5.1open | Z.ai | 93.3% | 2026-08-10 | independent | Epoch AI Benchmarking Hub |
| 16 | Grok 4.3 Beta | xAI | 93.3% | 2026-06-17 | independent | Epoch AI Benchmarking Hub |
| 17 | Qwen 3.6 Plus | Alibaba | 93.3% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 18 | Claude Opus 4.6 | Anthropic | 92.9% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 19 | Claude Opus 4.7 | Anthropic | 92.2% | 2026-08-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 20 | Grok 4.20 | xAI | 92.2% | 2026-07-13 | independent | Epoch AI Benchmarking Hub |
| 21 | Kimi K2.5open | Moonshot AI | 92.2% | 2026-02-02 | independent | Epoch AI Benchmarking Hub |
| 22 | Gemini 3 Pro | Google DeepMind | 91.4% | 2025-11-19 | independent | Epoch AI Benchmarking Hub |
| 23 | Qwen 3.6 Max (Preview) | Alibaba | 91.1% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 24 | Inkling-Smallopen | Thinking Machines | 90.0% | 2026-08-14 | independent | Epoch AI Benchmarking Hub |
| 25 | gpt-oss-120bopen | OpenAI | 88.9% | 2025-12-11 | independent | Epoch AI Benchmarking Hub |
| 26 | Inklingopen | Thinking Machines | 88.9% | 2026-08-05 | independent | Epoch AI Benchmarking Hub |
| 27 | Muse Spark | Meta AI | 88.9% | 2026-04-08 | independent | Epoch AI Benchmarking Hub |
| 28 | Gemini 3.5 Flash | Google DeepMind | 88.1% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 29 | GPT-5.6 Sol | OpenAI | 88.1% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 30 | Claude Sonnet 5 | Anthropic | 87.4% | 2026-08-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 31 | Qwen 3.5 Plus (hosted 397B-A17B) | Alibaba | 86.7% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 32 | Qwen3-235B-A22B (Jul 2025)open | Alibaba | 86.7% | 2025-12-10 | independent | Epoch AI Benchmarking Hub |
| 33 | Qwen3.7-Plus | Alibaba | 86.7% | 2026-08-07 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 34 | Kimi K3open | Moonshot AI | 86.5% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 35 | GPT-5.4 | OpenAI | 86.2% | 2026-07-15 | independent · 5 runs | Epoch AI Benchmarking Hub |
| 36 | Qwen3.5 397B-A17Bopen | Alibaba | 85.6% | 2026-08-07 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 37 | Gemini 3.6 Flash | Google DeepMind | 85.5% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 38 | GPT-5.2 | OpenAI | 85.4% | 2026-07-13 | independent · 5 runs | Epoch AI Benchmarking Hub |
| 39 | Qwen 3.5 Flash (hosted 35B-A3B) | Alibaba | 84.4% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 40 | Qwen 3.6 Flash | Alibaba | 84.4% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 41 | Gemini 2.5 Pro (Jun 2025) | Google DeepMind | 84.2% | 2025-11-16 | independent | Epoch AI Benchmarking Hub |
| 42 | Grok 4 | xAI | 84.0% | — | independent | Epoch AI Benchmarking Hub |
| 43 | GLM-4.7open | Z.ai | 83.3% | 2026-01-29 | independent | Epoch AI Benchmarking Hub |
| 44 | Kimi K2 Thinkingopen | Moonshot AI | 83.1% | 2025-11-13 | independent | Epoch AI Benchmarking Hub |
| 45 | Qwen3.7 Flash | Alibaba | 82.2% | 2026-08-07 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 46 | GPT-5.5 | OpenAI | 80.7% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 47 | GPT-5.6 Terra | OpenAI | 80.6% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 48 | GLM-5open | Z.ai | 80.0% | 2026-02-12 | independent | Epoch AI Benchmarking Hub |
| 49 | DeepSeek-V4-Proopen | DeepSeek | 79.6% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 50 | Qwen3.6 27Bopen | Alibaba | 78.9% | 2026-08-07 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 51 | Claude Sonnet 4.6 | Anthropic | 78.7% | 2026-08-06 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 52 | Qwen 3.6 35B-A3Bopen | Alibaba | 77.8% | 2026-08-07 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 53 | o3 | OpenAI | 76.1% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 54 | GPT-5 | OpenAI | 75.1% | 2026-07-20 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 55 | GPT-5 mini | OpenAI | 73.5% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 56 | Gemma 4 31B ITopen | Google DeepMind | 73.3% | 2026-08-06 | independent | Epoch AI Benchmarking Hub |
| 57 | Qwen3-Max | Alibaba | 73.3% | 2025-10-06 | independent | Epoch AI Benchmarking Hub |
| 58 | Claude Opus 4.5 | Anthropic | 71.9% | 2025-11-24 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 59 | Gemini 2.5 Flash | Google DeepMind | 71.9% | 2025-05-20 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 60 | MiniMax-M3open | MiniMax | 71.1% | 2026-08-10 | independent | Epoch AI Benchmarking Hub |
| 61 | o4-mini | OpenAI | 70.9% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 62 | Grok-3 mini | xAI | 70.0% | 2025-04-10 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 63 | GPT-5.1 | OpenAI | 69.0% | 2026-08-07 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 64 | DeepSeek-V3.2open | DeepSeek | 68.4% | 2026-07-16 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 65 | GPT-5.6 Luna | OpenAI | 68.3% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 66 | GPT-5.5 Instant | OpenAI | 68.1% | 2026-08-02 | independent | Epoch AI Benchmarking Hub |
| 67 | GPT-5.4 Nano | OpenAI | 67.8% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 68 | GPT-5.4 Mini | OpenAI | 67.6% | 2026-08-07 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 69 | o1 | OpenAI | 66.7% | 2026-07-15 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 70 | DeepSeek-R1 (May 2025)open | DeepSeek | 66.4% | 2025-05-29 | independent | Epoch AI Benchmarking Hub |
| 71 | Claude Sonnet 4.5 | Anthropic | 65.6% | 2025-10-28 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 72 | o3-mini | OpenAI | 61.8% | 2026-07-20 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 73 | Gemini 3.5 Flash-Lite | Google DeepMind | 60.7% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 74 | GPT-5 nano | OpenAI | 59.4% | 2026-07-13 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 75 | Claude Opus 4.1 | Anthropic | 57.8% | 2025-08-05 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 76 | Gemini 2.0 Flash Thinking | Google DeepMind | 57.8% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 77 | GLM-5.2open | Z.ai | 57.6% | 2026-08-10 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 78 | Claude Opus 4 | Anthropic | 55.6% | 2025-05-28 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 79 | Claude Sonnet 4 | Anthropic | 55.6% | 2025-05-23 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 80 | Grok 3 | xAI | 55.6% | 2025-04-10 | independent | Epoch AI Benchmarking Hub |
| 81 | Gemini 3.1 Flash-Lite | Google DeepMind | 54.1% | 2026-08-06 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 82 | gpt-oss-20bopen | OpenAI | 53.9% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 83 | DeepSeek-R1open | DeepSeek | 53.3% | 2025-02-26 | independent | Epoch AI Benchmarking Hub |
| 84 | DeepSeek-R1-Distill-Llama-70Bopen | DeepSeek | 51.4% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 85 | Claude Haiku 4.5 | Anthropic | 51.3% | 2025-10-22 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 86 | o1-mini | OpenAI | 45.8% | 2025-03-06 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 87 | Claude 3.7 Sonnet | Anthropic | 44.9% | 2025-03-13 | independent · 4 runs | Epoch AI Benchmarking Hub |
| 88 | GPT-4.1 mini | OpenAI | 44.7% | 2025-04-14 | independent | Epoch AI Benchmarking Hub |
| 89 | GPT-4.1 | OpenAI | 38.3% | 2025-04-14 | independent | Epoch AI Benchmarking Hub |
| 90 | DeepSeek-V3 (Mar 2025)open | DeepSeek | 37.8% | 2025-04-01 | independent | Epoch AI Benchmarking Hub |
| 91 | GPT-4.5 | OpenAI | 37.8% | 2025-02-28 | independent | Epoch AI Benchmarking Hub |
| 92 | Mistral Medium 3 | Mistral | 32.2% | 2025-05-07 | independent | Epoch AI Benchmarking Hub |
| 93 | Gemini 2.0 Flash | Google DeepMind | 31.1% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 94 | o1-preview | OpenAI | 31.1% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 95 | Magistral Small 1.0open | Mistral | 30.0% | 2026-08-06 | independent | Epoch AI Benchmarking Hub |
| 96 | GPT-4.1 nano | OpenAI | 28.9% | 2025-04-14 | independent | Epoch AI Benchmarking Hub |
| 97 | Gemma 3 27Bopen | Google DeepMind | 22.2% | 2026-08-06 | independent | Epoch AI Benchmarking Hub |
| 98 | Llama 4 Maverickopen | Meta AI | 20.6% | 2025-04-08 | independent | Epoch AI Benchmarking Hub |
| 99 | Qwen Plus | Alibaba | 17.8% | 2025-04-07 | independent | Epoch AI Benchmarking Hub |
| 100 | Qwen2.5-Max | Alibaba | 16.1% | 2025-04-01 | independent | Epoch AI Benchmarking Hub |
| 101 | DeepSeek-V3open | DeepSeek | 15.8% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 102 | Gemini 1.5 Pro | Google DeepMind | 14.9% | 2025-02-25 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 103 | Phi-4open | Microsoft | 13.8% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 104 | Grok-2 | xAI | 11.5% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 105 | Gemini 1.5 Flash | Google DeepMind | 10.1% | 2025-02-25 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 106 | Llama 3.1-405Bopen | Meta AI | 9.7% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 107 | Claude 3.5 Sonnet (October 2024) | Anthropic | 8.5% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 108 | Mistral Large 2open | Mistral | 8.1% | 2025-02-25 | independent · 2 runs | Epoch AI Benchmarking Hub |
| 109 | Qwen2.5-72Bopen | Alibaba | 8.1% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 110 | Llama 4 Scoutopen | Meta AI | 7.8% | 2025-04-08 | independent | Epoch AI Benchmarking Hub |
| 111 | Qwen2.5-32Bopen | Alibaba | 7.4% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 112 | GPT-4o mini | OpenAI | 6.9% | 2025-07-30 | independent | Epoch AI Benchmarking Hub |
| 113 | GPT-4 Turbo (Apr 2024) | OpenAI | 6.7% | 2025-02-27 | independent | Epoch AI Benchmarking Hub |
| 114 | Claude 3.5 Sonnet | Anthropic | 6.5% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 115 | GPT-4o | OpenAI | 6.3% | 2025-02-25 | independent · 3 runs | Epoch AI Benchmarking Hub |
| 116 | Qwen-Turbo | Alibaba | 6.1% | 2025-04-07 | independent | Epoch AI Benchmarking Hub |
| 117 | Mistral Small 3.1open | Mistral | 5.8% | 2025-03-18 | independent | Epoch AI Benchmarking Hub |
| 118 | Mistral Small 3open | Mistral | 5.3% | 2025-03-18 | independent | Epoch AI Benchmarking Hub |
| 119 | Llama 3.3 70Bopen | Meta AI | 5.1% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 120 | Claude 3 Opus | Anthropic | 4.7% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 121 | Gemini 1.5 Flash 8B | Google DeepMind | 4.6% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 122 | Tulu 3 (Tülu 3) 70Bopen | Allen Institute | 4.4% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 123 | Claude 3.5 Haiku | Anthropic | 4.3% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 124 | Llama 3-70Bopen | Meta AI | 4.3% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 125 | Llama 3.1-70Bopen | Meta AI | 3.6% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 126 | Llama 3.2 90Bopen | Meta AI | 2.6% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 127 | Claude 2 | Anthropic | 2.5% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 128 | Claude 3 Sonnet | Anthropic | 2.5% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 129 | Hermes 2 Theta Llama-3 70Bopen | Other | 2.5% | 2025-04-01 | independent | Epoch AI Benchmarking Hub |
| 130 | Llama 3.1-8Bopen | Meta AI | 2.5% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 131 | GPT-3.5 Turbo | OpenAI | 2.2% | 2026-08-07 | independent | Epoch AI Benchmarking Hub |
| 132 | Claude 2.1 | Anthropic | 1.9% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 133 | Mistral Large | Mistral | 1.9% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 134 | Claude 3 Haiku | Anthropic | 1.8% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 135 | Gemma 2 27Bopen | Google DeepMind | 1.4% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 136 | Gemini 1.0 Pro | Google DeepMind | 1.1% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 137 | GPT-4 (Jun 2023) | OpenAI | 1.1% | 2025-10-23 | independent | Epoch AI Benchmarking Hub |
| 138 | Llama 3-8Bopen | Meta AI | 0.8% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |
| 139 | Gemma 2 9Bopen | Google DeepMind | 0.6% | 2025-03-07 | independent | Epoch AI Benchmarking Hub |
| 140 | GPT-4 (Mar 2023) | OpenAI | 0.6% | 2025-10-22 | independent | Epoch AI Benchmarking Hub |
| 141 | Llama 2-70Bopen | Meta AI | 0.0% | 2025-02-25 | independent | Epoch AI Benchmarking Hub |