all leaderboards
vetto research

GulliBench

The official set is 50 multi-hop tasks, the hardest we found, each rendered in three substrates (table, docs, and SQL), for 150 tasks. 18 frontier models, a single MCP agent, five runs per task. Pass rate is the mean pass@1 (n=5): how often the model resisted both traps and derived the truth, so higher is less gullible.

150 tasks · 50 multi-hop × 3 substrates pass@1 · 5 trials/task 18 frontier models · single MCP agent updated August 2026
official board

Pass rate by model

Pooled across the three substrates or split by one. Whiskers are 95% bootstrap CIs; where two models sit inside each other's intervals, the board reads as a tie.

Substrate
breakdown
Pass rate by task & model
Substrate
fooled 0%100% derives from source