← all leaderboards
vetto research
GulliBench
The official set is 50 multi-hop tasks, the hardest we found, each rendered in three substrates (table, docs, and SQL), for 150 tasks. 18 frontier models, a single MCP agent, five runs per task. Pass rate is the mean pass@1 (n=5): how often the model resisted both traps and derived the truth, so higher is less gullible.
official board
Pass rate by model
Pooled across the three substrates or split by one. Whiskers are 95% bootstrap CIs; where two models sit inside each other's intervals, the board reads as a tie.
Substrate
breakdown
Pass rate by task & model
Substrate
fooled 0%100% derives from source