← All benchmarks

EBR-bench

Reasoning unit: % independent 18 models scored not in composite

Evidence-based reasoning tasks that require weighing conflicting sources before answering.

What it measures

Judgement over retrieval: what to believe when the evidence disagrees.

How to read it

Newer benchmark with fewer models evaluated so far.

Source

Epoch AI Benchmarking Hub

Full ranking

#ModelLabScoreMeasuredMethodSource
1 Claude Opus 5 Anthropic 50.0% 2026-08-07 independent Epoch AI Benchmarking Hub
2 Claude Fable 5 Anthropic 39.5% 2026-07-17 independent Epoch AI Benchmarking Hub
3 GPT-5.6 Sol OpenAI 39.0% 2026-08-07 independent Epoch AI Benchmarking Hub
4 GPT-5.5 OpenAI 34.3% 2026-07-27 independent Epoch AI Benchmarking Hub
5 Claude Opus 4.8 Anthropic 28.6% 2026-08-07 independent Epoch AI Benchmarking Hub
6 GPT-5.4 OpenAI 25.4% 2026-06-25 independent Epoch AI Benchmarking Hub
7 GPT-5.2 OpenAI 23.0% 2026-06-26 independent Epoch AI Benchmarking Hub
8 Claude Opus 4.7 Anthropic 19.0% 2026-06-30 independent Epoch AI Benchmarking Hub
9 Claude Opus 4.5 Anthropic 14.3% 2026-06-25 independent Epoch AI Benchmarking Hub
10 Gemini 3.1 Pro Google DeepMind 14.3% 2026-06-25 independent Epoch AI Benchmarking Hub
11 Claude Opus 4.6 Anthropic 12.7% 2026-06-29 independent Epoch AI Benchmarking Hub
12 GPT-5 OpenAI 12.7% 2026-06-29 independent Epoch AI Benchmarking Hub
13 GLM-5.2open Z.ai 9.5% 2026-06-29 independent Epoch AI Benchmarking Hub
14 Qwen3.7-Max Alibaba 9.5% 2026-06-25 independent Epoch AI Benchmarking Hub
15 Claude Opus 4.1 Anthropic 7.9% 2026-06-25 independent Epoch AI Benchmarking Hub
16 Gemini 3.5 Flash Google DeepMind 4.8% 2026-06-25 independent Epoch AI Benchmarking Hub
17 Claude Sonnet 4.5 Anthropic 2.4% 2026-06-25 independent Epoch AI Benchmarking Hub
18 Kimi K2.6open Moonshot AI 2.4% 2026-06-25 independent Epoch AI Benchmarking Hub