Plot Twist Bench
Twelve models on 48 multiple-choice questions over 26 recordings filmed to order, each built so the usual thing is not what happened. Every question was asked once per ordering of its options, so the headline is pass@1: the average across those rotations. Grading is a plain string comparison against a human-annotated golden answer.
pass@1 by model
Ranked by pass@1, the average across every ordering of a question's options. Whiskers are 95% intervals clustered on the 26 recordings. The five testers answered each question once too, so the human baseline compares directly to these bars.
A question answered right under three of four orderings scores 0.75. Since attempts differ only in option order, that average is the model's chance of a right answer on a single try. Whiskers are 95% intervals clustered on the 26 recordings.
Right under every ordering
Circular accuracy counts a question only if the model is right under every one of its orderings. It carries its own cluster-robust interval, and the two are never mixed.
No human row here: the 85% baseline is a single-pass number, so no model was asked to clear it under this measure. Exported 26 August 2026.