all leaderboards
vetto research

GulliBench

The official set is 50 multi-hop tasks, the hardest we found, each rendered in three substrates (table, docs, and SQL), for 150 tasks. 18 frontier models, a single MCP agent, five runs per task. Pass rate is the mean pass@1 (n=5): how often the model resisted both traps and derived the truth, so higher is less gullible.

150 tasks · 50 multi-hop × 3 substrates pass@1 · 5 trials/task 18 frontier models · single MCP agent updated August 2026
read the research post
official board

Pass rate by model

Pooled across the three substrates or split by one. Whiskers are 95% bootstrap CIs; where two models sit inside each other's intervals, the board reads as a tie.

Substrate
# model 0–100%
1 Opus 5 Anthropic 49.3%
2 Fable 5 Anthropic 48.1%
3 Muse Spark 1.2 Meta 42.0%
4 Gemini 3.1 Pro Google DeepMind 22.5%
5 Kimi K3 Moonshot AI 17.6%
6 Grok 4.6 xAI 16.3%
7 Opus 4.8 Anthropic 15.9%
8 DeepSeek V4 Flash DeepSeek 15.7%
9 GLM 5.2 Z.ai 15.1%
10 DeepSeek V4 Pro DeepSeek 11.5%
11 Sonnet 5 Anthropic 8.3%
12 GPT-5.5 OpenAI 6.5%
13 Gemini 3.5 Flash Google DeepMind 3.1%
14 Gemini 3.7 Flash Google DeepMind 2.8%
15 Gemini 3.6 Flash Google DeepMind 2.7%
16 GPT-5.6 Sol OpenAI 2.4%
17 GPT-5.6 Terra OpenAI 2.1%
18 GPT-5.6 Luna OpenAI 2.1%

Official set · 50 multi-hop tasks · all substrates · pass % (mean pass@1, n=5) · whisker = 95% CI

breakdown

Pass rate by task & model

The share of each model's five trials on each task that derived the truth. Hover a cell for the counts.

Substrate
model ↓ · task → 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50
Muse Spark 1.2 100 80 100 100 80 100 100 100 100 100 100 100 60 100 20 60 60 60 60 100 40 40 60 60 60 100 40 40 40 0 80 60 40 60 0 0 80 20 40 0 80 20 60 60 0 0 0 0 0 0
Opus 5 0 100 100 0 100 80 100 100 0 0 0 80 100 80 80 100 100 40 100 40 100 60 100 100 0 60 60 60 20 100 0 60 80 0 100 100 20 40 80 40 0 0 0 0 0 0 0 0 0 0
Fable 5 40 100 40 40 100 0 100 80 0 0 0 60 60 100 80 80 60 80 60 20 80 100 0 40 100 0 60 0 100 60 40 20 20 60 40 40 0 0 0 60 0 40 40 0 40 40 0 20 0 0
Gemini 3.1 Pro 100 20 80 100 80 100 20 60 100 100 100 0 80 0 20 20 0 0 0 20 0 0 20 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 20 20 0 0 0 0 0 20 0 0 0
Grok 4.6 80 80 100 80 100 40 40 20 60 100 100 40 0 0 20 0 0 0 0 0 0 0 0 0 0 0 0 20 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
Kimi K3 80 0 100 60 80 80 60 60 40 40 0 0 40 0 0 20 20 0 0 20 0 0 20 0 20 20 0 0 0 0 20 0 0 20 0 0 20 20 0 0 0 20 0 0 0 0 0 0 0 0
Opus 4.8 0 100 100 0 100 0 100 80 0 0 0 100 20 40 100 0 0 80 0 0 0 0 0 0 0 0 20 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
GLM 5.2 100 40 60 40 60 60 40 60 60 80 40 0 20 40 20 0 0 0 0 0 0 20 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
DeepSeek V4 Flash 100 80 0 0 20 60 60 0 20 80 40 60 0 0 0 0 20 0 0 20 0 0 0 0 0 0 0 40 0 0 0 0 0 0 0 0 0 20 0 0 0 20 0 0 0 0 0 0 0 0
DeepSeek V4 Pro 100 80 20 80 20 80 0 0 80 20 60 0 0 0 20 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
Sonnet 5 0 100 60 40 0 20 80 20 20 0 0 20 20 0 0 0 0 0 0 0 0 0 0 0 20 0 0 0 0 0 0 0 0 0 0 0 0 20 0 0 0 0 0 0 0 0 0 0 0 0
GPT-5.5 80 0 20 100 0 0 0 0 0 0 40 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
Gemini 3.5 Flash 20 20 0 80 0 40 0 0 20 20 20 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
GPT-5.6 Luna 0 20 0 0 0 40 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
Gemini 3.6 Flash 20 0 0 20 0 0 0 0 20 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
Gemini 3.7 Flash 40 0 0 0 0 0 0 0 20 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
GPT-5.6 Sol 20 0 20 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
GPT-5.6 Terra 20 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
fooled 0% 100% derives from source