all leaderboards
vetto research

Plot Twist Bench

Twelve models on 48 multiple-choice questions over 26 recordings filmed to order, each built so the usual thing is not what happened. Every question was asked once per ordering of its options, so the headline is pass@1: the average across those rotations. Grading is a plain string comparison against a human-annotated golden answer.

48 questions · 26 recordings pass@1 · every ordering 12 frontier models human baseline 85% updated August 2026
read the research post
official board

pass@1 by model

Ranked by pass@1, the average across every ordering of a question's options. Whiskers are 95% intervals clustered on the 26 recordings. The five testers answered each question once too, so the human baseline compares directly to these bars.

# model 0–100%
1 Human baseline 5 testers · 84–88 85%
2 GPT-5.6 Sol OpenAI 75%
3 GPT-5.5 OpenAI 75%
4 Kimi K3 Moonshot AI 73%
5 Gemini 3.7 Flash Google DeepMind 72%
6 GPT-5.6 Terra OpenAI 71%
7 Gemini 3.6 Flash Google DeepMind 70%
8 Claude Opus 5 Anthropic 63%
9 Qwen 3.7 Plus Alibaba 58%
10 Grok 4.3 xAI 57%
11 Gemini 3.1 Pro Google DeepMind 56%
12 Claude Sonnet 4.6 Anthropic 49%
13 Claude Opus 4.8 Anthropic 41%

A question answered right under three of four orderings scores 0.75. Since attempts differ only in option order, that average is the model's chance of a right answer on a single try. Whiskers are 95% intervals clustered on the 26 recordings.

strict companion

Right under every ordering

Circular accuracy counts a question only if the model is right under every one of its orderings. It carries its own cluster-robust interval, and the two are never mixed.

# model 0–100%
1 GPT-5.6 Sol OpenAI 67%
2 Kimi K3 Moonshot AI 67%
3 GPT-5.5 OpenAI 65%
4 Gemini 3.6 Flash Google DeepMind 62%
5 Gemini 3.7 Flash Google DeepMind 60%
6 GPT-5.6 Terra OpenAI 54%
7 Claude Opus 5 Anthropic 50%
8 Qwen 3.7 Plus Alibaba 46%
9 Grok 4.3 xAI 42%
10 Claude Sonnet 4.6 Anthropic 40%
11 Claude Opus 4.8 Anthropic 27%
12 Gemini 3.1 Pro Google DeepMind 25%

No human row here: the 85% baseline is a single-pass number, so no model was asked to clear it under this measure. Exported 26 August 2026.