Plot Twist Bench
Do frontier models actually understand video? Plot Twist Bench is a video perception benchmark of 48 curated multiple-choice questions, made by Vetto.
01The question every model missed
Three small clusters of objects on a table, filmed from directly above, and a hand that moves a couple of items between them. Twenty-seven seconds, one fixed camera, nothing cut and nothing hidden.
We asked twelve frontier models one question about it: which cluster reveals an object that was not perceivable at the start?
The answer is the rightmost. Three buttons were placed there and only two could ever be seen, because the third was resting on top of one of the others. When the hand lifted it away, the button underneath became a separate object for the first time.
None of the twelve got it. Nine answered leftmost, the cluster the button arrives in. The other three, all of them Gemini models, answered that nothing was revealed at all.
That question is one of the 48 in the set. To understand why Gemini 3.1 Pro got it wrong, we put two more questions to it once the graded run was over.
Why is this question hard for you?
I didn’t actually find it hard to answer! The video provides a clear, top-down view of all the objects from the very first frame. Because all the peanuts and the two blue beads are fully visible in their clusters before any movement happens, it is straightforward to observe that when the hand transfers the peanut and the blue bead, no hidden objects are uncovered or revealed in the process. Everything was already in plain sight.
But where was the blue bead that ends up in the leftmost cluster?
The blue bead that ends up in the leftmost cluster was originally located in the rightmost cluster. At the 0:13 mark, the hand picks it up from the rightmost group and moves it over to the leftmost group. It was visible in that rightmost cluster from the very beginning of the video.
The model tracks the moved object correctly, so it names it in the second answer. However, it never checks that fact against the question, and doesn’t notice the gap.
This is the core problem this bench measures. The model can name the objects and describe the scene correctly, but it does not track what actually happened on screen: it settles on a description early and defends it, rather than checking it against what’s on screen. That gap is why it gives a confident wrong answer.
Recognition is not perception
A vision-language model predicts the most likely continuation of what it’s given. Most of the time that looks like watching, because the likeliest description of a scene is usually the true one.
Two abilities blend together here, and this set separates them. Recognition names the objects and the type of scene; it can be answered from what usually happens. Perception tracks what actually happened on screen; it cannot be answered that way.
Telling them apart needs footage built for the purpose, so we filmed our own: thirteen scripts, twenty-six takes, never published online, each built around one twist decided before filming. That is Plot Twist Bench: 48 questions over 26 recordings.
02The leaderboard
How the score works. Twelve models, 48 questions. Each model answered every question once per ordering of its options, so a four-option question was asked four times, with the options rotated each time. The headline number is pass@1: the average across those rotations. A question answered right under three of four orderings scores 0.75. Since attempts differ only in option order, that average is the model’s chance of a right answer on a single try. Circular accuracy (CircularEval) is the stricter measure: it counts a question only if the model is right under every ordering. Grading is a plain string comparison against a human annotated golden answer.
all twelve models
pass@1: 75% at the top, people at 85%
-
1
GPT-5.6 Sol
OpenAI
75
% -
2
GPT-5.5
OpenAI
75
% -
3
Kimi K3
Moonshot AI
73
% -
4
Gemini 3.7 Flash
Google
72
% -
5
GPT-5.6 Terra
OpenAI
71
% -
6
Gemini 3.6 Flash
Google
70
% -
7
Claude Opus 5
Anthropic
63
% -
8
Qwen 3.7 Plus
Alibaba
58
% -
9
Grok 4.3
xAI
57
% -
10
Gemini 3.1 Pro
Google
56
% -
11
Claude Sonnet 4.6
Anthropic
49
% -
12
Claude Opus 4.8
Anthropic
41
%
The five testers also answered each question once, so the dashed 85% line compares directly to these bars.
What the scores show. The field spans a wide range, from 75% at the top to 41% at the bottom, and every model falls short of the human baseline of 85%. The two scores diverge for two reasons: missing the answer outright, or getting it right once and failing to repeat it. That second failure is uneven, costing steady models little and unstable ones a lot. Seeing the scene once is within roughly 10 points of the 85% human baseline; holding onto what was seen is where the field comes apart.
Those pass@1 and circular-accuracy numbers carry real uncertainty, and the interval on each row is deliberately conservative: because a model’s answers on the same clip move together (within-clip answer correlation, ICC, is roughly 0.2), we cluster uncertainty by video rather than by question, widening the bars but keeping them honest. How the intervals are computed →
That gap between seeing something once and holding onto it comes down to consistency, which breaks into two shapes, and recent work argues for keeping them apart (Romero-Alvarado et al., 2026). A model can be consistently wrong, repeating the same mistake every reordering, or inconsistently right, slipping between correct and incorrect readings. The repeatability cost itself ranges from about seven points for the steadiest models to as much as thirty-one for the most unstable. Claude Sonnet 4.6 and GPT-5.5 are both 71% consistent, yet Sonnet trails by 25 points on accuracy: one holds firmly to the wrong answer, the other doesn’t waver at all. Consistency per model →
All of that instability sits below the same human ceiling, which is worth grounding in specifics: five testers who had never seen the questions answered them cold and still cleared it, scoring between 84% and 88% and averaging 85%. This is the standard the model gaps are measured against. How the baseline was measured →
What each model actually saw. Some of that score gap traces back to input, not just ability: each row shows a model on its best available video path, not a common setup for all twelve: the three Gemini models took a video file directly, the rest ran on stills with a transcript and spectrogram. What each model received →
Each model ran on its best available setup by design, and the baseline comparison still holds up to a point: three models, GPT-5.6 Sol, GPT-5.5 and Kimi K3, fall within the human baseline’s margin of error on a single pass, but no model matches humans once the same question is asked with reshuffled answer options.
What this extends
DeepMind’s Perception Test (Pătrăucean et al., NeurIPS 2023) is the work that convinced us to film rather than scrape, and it’s the closest predecessor to this one. We took the premise from it; what’s changed since is the score. At publication the best models managed 46.2% against a 91.4% human baseline. By the challenge write-up in January 2026, the top entries had reached 83.7% on the original multiple-choice set.
Those entries were ensembles built on Seed 1.6 Vision, Qwen-VL-Max and Qwen-2.5VL. Every model on our board is a generation newer and stronger. The best of the twelve averages 75% here, and holds 67% when the answer has to survive every reordering of its options.
Price does not buy sight
The figures are metered spend over the same 26 recordings, every model called on every one, normalised to what one clean run of a question costs. The quality axis is the headline pass@1, and the units are defined in the appendix.
cost–quality frontier · pass@1
Nine of the twelve sit off the frontier
Hover a point, or tab to it, for that model’s full figures.Tap any point for that model’s name and full figures.
Three models hold the frontier, and two of them are one model a version apart: Gemini 3.6 Flash at $0.021 a question for 70%, Gemini 3.7 Flash at $0.042 for 72%, and GPT-5.6 Sol at $0.426 for 75%. The other nine sit off it, each matched or beaten on both axes at once, so no budget makes them the right pick on this set.
Sol’s 3-point lead over Gemini 3.7 Flash sits well inside what this set can resolve; its 10× price does not. GPT-5.5 matches Sol’s 75% at a sixth more spend. Kimi K3 sits two points under Sol at 2.5× the price. The sharpest case sits inside one family: Gemini 3.6 Flash is 62% cheaper per question and 14 points better than Gemini 3.1 Pro.
03Task design
Each card below opens one recording: the clip, the question as the model received it, the shooting script behind it and every model’s grade. Only three appear here, because this section is a deep dive rather than a measurement: each card is picked to show one failure clearly, not to stand in for the whole set. Two of the three were failed by every model on the board; the third was passed by just one of twelve.
That selection is why the mechanism named in each card is our own reading rather than a measured rate: nothing here says how often the failure happens across the set. For that, the failure landscape has the rates, across all 48 questions and all twelve models.
04The failure landscape
Every mode, every model
Every question is tagged by hand, before any model runs, with the failure modes it is meant to expose, and those tags were audited again after the runs are graded. That gives a rate per model per mode: the share of that mode’s questions the model passed, where a pass is the strict one, right under every ordering. All 20 modes this set exercises, ordered by how much evidence each carries.
20 failure modes · 12 models · 240 rates
Where Gemini 3.1 Pro loses, and to whom
| Failure mode | Questions | Gemini 3.1 ProTarget | Other eleven modelsField | DeltaΔ |
|---|---|---|---|---|
| 14 | 7% |
49% | -42 | |
|
What it measures Judging relative quantity or magnitude, which is more, less, or larger, including estimating a count in a dense cluster rather than enumerating it exactly. All twelve models on this mode
|
||||
| 12 | 8% |
52% | -44 | |
|
What it measures Perceiving, matching, and telling apart objects by colour, size, material, shape, or other visible attributes, including near-identical items. All twelve models on this mode
|
||||
| 10 | 50% |
56% | -6 | |
|
What it measures Recalling what happened earlier in the clip, independent of where it falls in the sequence. All twelve models on this mode
|
||||
| 7 | 29% |
64% | -35 | |
|
What it measures The correct order of events, at the narrative level or, in stress-test form, the raw frame level. All twelve models on this mode
|
||||
| 6 | 17% |
33% | -16 | |
|
What it measures An object still exists, and stays where it was, while hidden. All twelve models on this mode
|
||||
| 6 | 83% |
70% | +13 | |
|
What it measures Counting discrete objects, actions, or events exactly. All twelve models on this mode
|
||||
| 6 | 33% |
65% | -32 | |
|
What it measures Where things are relative to each other; what is inside or contains what. All twelve models on this mode
|
||||
| 5 | 40% |
45% | -5 | |
|
What it measures Ignoring an irrelevant object or action, whether it was deliberately planted to mislead or is simply salient-but-beside-the-point. All twelve models on this mode
|
||||
| 4 | 25% |
61% | -36 | |
|
What it measures Following two or more simultaneous, equally relevant action streams at once (e.g. two pans cooking). All twelve models on this mode
|
||||
| 4 | 25% |
50% | -25 | |
|
What it measures Naming objects and their parts. All twelve models on this mode
|
||||
| 3 | 0% |
58% | -58 | |
|
What it measures Noticing that something changed between two moments, a model can name both states correctly in isolation and still miss that a change happened. All twelve models on this mode
|
||||
| 3 | 0% |
52% | -52 | |
|
What it measures Inferring an interaction between objects that happens partly out of view, behind an occluder. All twelve models on this mode
|
||||
| 3 | 0% |
24% | -24 | |
|
What it measures Pinning when an event occurs. All twelve models on this mode
|
||||
| 2 | 0% |
9% | -9 | |
|
What it measures Mapping left/right and handedness from the camera's frame onto a subject's own frame (e.g. "their right hand" vs. "screen right"). All twelve models on this mode
|
||||
| 2 | 0% |
45% | -45 | |
|
What it measures Distinguishing an action performed to deceive the observer, mimed, staged as a decoy, deliberately misdirecting, or falsely presented as having completed a task, from one that's genuinely performed. All twelve models on this mode
|
||||
| 1 | 0% |
18% | -18 | |
|
What it measures Reasoning through a chain of events if one condition had changed. All twelve models on this mode
|
||||
| 1 | 0% |
55% | -55 | |
|
What it measures Detecting a violation in a repeating pattern. All twelve models on this mode
|
||||
| 1 | 100% |
73% | +27 | |
|
What it measures Identifying the location, room, or scene. All twelve models on this mode
|
||||
| 1 | 0% |
64% | -64 | |
|
What it measures Forecasting a future physical outcome or imminent event that isn't specifically a support/stability or solidity/collision case (e.g. trajectory, timing, momentum). All twelve models on this mode
|
||||
| 1 | 0% |
27% | -27 | |
|
What it measures An object's or scene's state at a single moment (open/closed, on/off, cooked/raw), including whether a stated goal or task was actually reached. All twelve models on this mode
|
||||
| Failure mode | G3.1 Pro |
Gap | Field | Sol | Kimi K3 | GPT-5.5 | G3.6 Flash | G3.7 Flash | Terra | Opus 5 | Qwen 3.7 | Grok 4.3 | Sonnet 4.6 | Opus 4.8 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Quantity comparison & estimation |
7% | -42 | 49% | 79% | 86% | 79% | 43% | 36% | 57% | 57% | 36% | 36% | 29% | 7% |
| Object attribute discrimination |
8% | -44 | 52% | 58% | 58% | 58% | 75% | 75% | 50% | 50% | 42% | 42% | 42% | 25% |
| Event recall |
50% | -6 | 56% | 60% | 60% | 50% | 70% | 80% | 40% | 50% | 50% | 50% | 70% | 40% |
| Event ordering |
29% | -35 | 64% | 71% | 57% | 71% | 86% | 71% | 71% | 71% | 57% | 57% | 43% | 43% |
| Object permanence |
17% | -16 | 33% | 50% | 33% | 67% | 50% | 33% | 33% | 17% | 17% | 33% | 17% | 17% |
| Object/action/event counting |
83% | +13 | 70% | 83% | 100% | 83% | 67% | 83% | 67% | 67% | 67% | 50% | 33% | 67% |
| Spatial relations & containment |
33% | -32 | 65% | 67% | 67% | 67% | 83% | 67% | 50% | 67% | 50% | 83% | 67% | 50% |
| Distractor & saliency resistance |
40% | -5 | 45% | 40% | 60% | 40% | 80% | 80% | 40% | 40% | 40% | 40% | 20% | 20% |
| Multi-stream parallel tracking |
25% | -36 | 61% | 75% | 75% | 75% | 75% | 75% | 50% | 25% | 75% | 50% | 75% | 25% |
| Object & part recognition |
25% | -25 | 50% | 50% | 50% | 50% | 75% | 75% | 50% | 25% | 25% | 50% | 75% | 25% |
| Change detection |
0% | -58 | 58% | 67% | 67% | 67% | 67% | 67% | 67% | 33% | 67% | 33% | 67% | 33% |
| Motion & occluded interactions |
0% | -52 | 52% | 67% | 33% | 100% | 67% | 33% | 67% | 33% | 33% | 67% | 33% | 33% |
| Temporal localization precision |
0% | -24 | 24% | 33% | 33% | 33% | 33% | 33% | 33% | 0% | 33% | 0% | 33% | 0% |
| Chirality / egocentric reference |
0% | -9 | 9% | 50% | 50% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
| Deceptive vs. genuine action |
0% | -45 | 45% | 100% | 0% | 100% | 50% | 50% | 100% | 50% | 0% | 0% | 0% | 50% |
| Counterfactual chain reasoning |
0% | -18 | 18% | 0% | 0% | 0% | 100% | 100% | 0% | 0% | 0% | 0% | 0% | 0% |
| Pattern breaking |
0% | -55 | 55% | 100% | 100% | 100% | 100% | 0% | 100% | 100% | 0% | 0% | 0% | 0% |
| Place recognition |
100% | +27 | 73% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 0% | 0% | 0% |
| Predictive dynamics |
0% | -64 | 64% | 100% | 0% | 100% | 100% | 100% | 100% | 100% | 0% | 0% | 0% | 100% |
| State & goal-state recognition |
0% | -27 | 27% | 0% | 0% | 0% | 100% | 100% | 0% | 0% | 100% | 0% | 0% | 0% |
05How the set was built
Every recording starts as a written scenario that names the failure it’s meant to provoke; the scene is then built to trigger it. Each scenario targets one twist aimed at one shortcut, written down before filming starts. Here is one of them, for a scene about making a cup of tea:
The real task (making a cup of tea) is deliberately left INCOMPLETE in an adversarial way: the person pours hot water into the mug but NEVER puts the teabag in. […] Keep the teabag visible on the counter the whole time so the omission is verifiable from the frame, not just inferred. If the teabag ever enters the mug, retake.
A take that breaks its scenario isn’t fixed in editing. It’s reshot. The ground truth is written against the take actually filmed, not the script that asked for it.
On fairness: hard for the right reason
A low score is only interesting if the difficulty sits in perceiving the video, rather than in a confusing prompt, a wrong expected answer, over-demanding grading, or a question answerable from the text alone, any of which would mean the score measures our instrument rather than the model.
Our own pipeline runs into that conflict directly: a question gets harder by iterating against a model until it misses, but that’s also the fastest way to create difficulty that has nothing to do with the video, so failing a model never earns a question its place. Instead, the set is cut from the top of a ranking that never looks at model performance, led by a Q-Score: a 0 to 100 rating of the question purely as a piece of writing, judged from its text alone, for whether it’s unambiguous, stands on its own, and can be graded. A question with two defensible answers scores zero, not something in between.
The check that settles it is the human baseline already on the board: five people with no hand in writing the questions answered them correctly 85% of the time, and a confusing question, a wrong expected answer, or one that can’t be answered from the video wouldn’t have scored that well. So whatever the models are running into here is in the footage, not the instrument. How the baseline was measured →
seven stages · nine acts
How a question gets into the set
Act 1 of 9A scenario is written against a named failure mode, built around one twist that defeats the shortcut, and it ships with its first questions, tagged before a frame exists.
Act 2 of 9A person films the take. Every recording is made to order, and a take that breaks the scenario is refilmed rather than patched.
Act 3 of 9The question, the options and the answer are authored by hand against the take that was actually filmed.
Act 4 of 9The question is sharpened against a reference model until the model reliably misses it. What earns a rewrite is the reasoning behind the answer, not the verdict alone.
Act 5 of 9The first gate is a Q-Score, a 0 to 100 rating of the question as a piece of writing: is it unambiguous, self-contained, and gradable. It is scored from the question text alone, so it never sees a model answer, and a question with two defensible answers scores zero rather than averaging out to respectable.
Act 6 of 9Every question is also answered from its text alone, with no video. A confident correct blind answer counts against the question.
Act 7 of 9The grading itself is audited for scope and balance. This is the gate that undoes act 4: difficulty that came from our wording rather than from the video is cut here, not counted as a win.
Act 8 of 9A second person reviews the recording and the question, and has to arrive at the same answer independently. Only approved takes continue.
Act 9 of 9What survives is ranked on question quality, and the set is cut from the top: 48 questions over 26 recordings.
What that pipeline cost in human hours
Every question carries the work of at least two people, a median of 2.9, because the person who hardens a question is never the person who audits it. No question in the set is one annotator’s judgment of their own work.
06The takeaway
The models on this board beat most humans at mathematics and competition coding. Show them a hand moving buttons around a table for twenty-seven seconds, and the strongest of them leads with 75% pass@1, trailing five ordinary people by ten points.
Three models get close to how people score on average: GPT-5.6 Sol and GPT-5.5 hit 75%, Kimi K3 hits 73%, against the testers’ 85%. That’s close enough that, on just 48 questions, we can’t tell them apart from a single person taking the test once. But not one of them beats even the worst human tester, who still scored 84%. Then just shuffle the order of the answer choices, and the picture falls apart. Every model drops, some by as little as 7 points, one by a full 31. After that shuffle, not a single model comes within 18 points of how people do. The reason is simple: a model settles on a description of the scene early and defends that description rather than checking it against the footage, so one answer only tells you what story it committed to, not whether that story is true.
Until then, ask a video model twice.
07Caveats
What these numbers cannot tell you, volunteered rather than waited for.
- A row scores a transport and a model together, not the weights alone. Only the three Gemini rows ran on a first-party video path. For seven of the other nine, stills are what the model ships: GPT, Claude and Grok take no video on any API we could call. The last two are a different case. Kimi K3 and Qwen 3.7 Plus both take video on their own APIs and neither got it here, so read those two rows as image-vision scores.
- Read the ordering with the intervals in view. One question is worth 2.1 points. The interval shown is the conservative one, clustered on the 26 recordings rather than the 48 questions. Rows within about ±12 of each other on the headline metric, or ±17 on the strict one, are too close for this set to rank: the middle of the table reads as a band, not an order. The bottom is separated on both readings.
- Selection is quality-gated, not difficulty-gated. These questions survived a fairness audit; they are not the ones any model fails hardest, so they are not comparable to our difficulty-gated sets.
- The three sample tasks are a deep dive, not a sample. Each was chosen because it shows one failure clearly, so they cannot say how often that failure happens.
- Cost is observed spend normalised to one clean run, not a rate card. Raw request counts differ widely across models after retries and reruns, so each model’s spend is divided by its raw requests and multiplied by the calls one clean run needs. The three Gemini rows are metered from Google’s own billing rather than the gateway.
What’s next
More questions, and more recordings. This is the top priority. At 48 questions over 26 recordings, the interval is wide enough that much of the board reads as a band, not a ranking. The video is the sampling unit, so a new recording narrows that interval faster than another question on a clip we already have. The next round grows the corpus first, and the question count with it.
Fixed-Context beside the native harness. The same corpus already runs a second way: every model, including Gemini, gets identical stills at 1 fps and the same transcript. Two numbers per model on the same 48 questions is the only clean way to see how much of each gap is transport rather than perception. Until it ships, the Gemini-last result can’t be called a perception finding on its own.
More evidence per failure mode, and a wider human panel. Twelve of the 20 modes carry four questions or fewer, too thin to rank. The baseline also rests on just five testers. So the next round widens the panel and gives each mode its own human number, not just the set as a whole. Alongside that, we’ll mark by hand the exact frames each answer lives in, turning “the model missed it” into “the model missed it in these 1.8 seconds”.
Longer videos, and more original scripts. Both push the corpus further from the distribution these models trained on.
08Appendix
Everything here is detail the sections above lean on without stopping to explain.
The harness each model got
The aim was to give every model the best video path we could actually reach, rather than flattening all twelve onto a lowest common denominator. In practice that meant meeting each API at its own highest capability.
The three Gemini models were called first-party, through Google’s own Files API, so they got the recording itself as a video file with its native audio track. The other nine were called through Vercel AI Gateway, which is what “stills” in the table below means: the gateway carries images but no vendor video format, so each of those models got still frames cut at its own documented maximum, from 1024px up to 2576px, plus a transcript of the audio and a spectrogram so that nothing in the soundtrack was lost to them.
Kimi K3 and Qwen 3.7 Plus can both take a video file on their own APIs. Neither got that here, and the two rows arrived there differently. Kimi was never sent the file: we cannot reach Moonshot’s video upload through the gateway we use, so from the first request Kimi saw stills we cut from the recording, at 1920px. Qwen was sent the file on all 26 recordings and the gateway refused it all 26 times, so Qwen also saw stills, at 1024px. Either way what arrived was sampled frames, a whisper-1 transcript and a spectrogram rather than video, which is why those two read as image-vision scores.
The other seven stills rows are not fallbacks. GPT, Claude and Grok take no video input on any API we could call, so stills are what those models actually ship.
| Model | Reached through | What it received | Audio |
|---|---|---|---|
| Gemini 3.1 Pro | Google, first-party | The recording as a video file, via the Files API, decoded by Google | Native audio track |
| Gemini 3.6 Flash | Google, first-party | The recording as a video file, via the Files API, decoded by Google | Native audio track |
| Gemini 3.7 Flash | Google, first-party | The recording as a video file, via the Files API, decoded by Google | Native audio track |
| GPT-5.6 Sol | Vercel AI Gateway | 1536px stills at the API’s highest detail setting | Transcript and spectrogram |
| GPT-5.6 Terra | Vercel AI Gateway | 1536px stills, same profile | Transcript and spectrogram |
| GPT-5.5 | Vercel AI Gateway | 1536px stills, same profile | Transcript and spectrogram |
| Kimi K3 | Vercel AI Gateway | 1920px stills, up to 180 of them; its video upload is not reachable through the gateway, so the file was never sent | Transcript and spectrogram |
| Claude Opus 5 | Vercel AI Gateway | 2576px stills, the high-resolution tier | Transcript and spectrogram |
| Claude Opus 4.8 | Vercel AI Gateway | 2576px stills, the high-resolution tier | Transcript and spectrogram |
| Claude Sonnet 4.6 | Vercel AI Gateway | 1568px stills, its no-resize threshold | Transcript and spectrogram |
| Grok 4.3 | Vercel AI Gateway | 1024px stills, up to 180 of them | Transcript and spectrogram |
| Qwen 3.7 Plus | Vercel AI Gateway | Requested as a video file, rejected on all 26, so 1024px stills | Transcript and spectrogram |
Resolution and image count follow each provider’s own documented limits. Frame density is the one thing we hold constant rather than maximise, at a target of one frame a second, so that the board compares models rather than frame rates. Stills are sampled uniformly with no keyframe selection, and past a model’s image budget a long clip is tiled into 2×2 grids rather than thinned, so “47 frames” is a mean over emitted images, from 17 to 159, not 47 instants of footage. Grading is identical on every row: temperature 0 except Kimi K3, which its provider fixes at 1.0, so that row’s circular grading is the one that is stochastic.
What each column on the leaderboard means
How the 95% interval is computed
Forty-eight questions is not forty-eight independent trials. They come from 26 recordings, and questions asked about the same clip are not independent of each other: if a model misreads the scene, it tends to miss every question on it. Treating them as independent would make the intervals look tighter than the evidence supports, so the video is the sampling unit and the interval is clustered on it.
How much that costs a given row depends on how strongly that model’s outcomes clump by clip, which is what the intra-cluster correlation (ICC) reports. Near zero means a clip tells you little about the next question on it, and per-question counting was roughly right. High and positive means the model passes or fails a recording more or less as a whole, so its effective sample is closer to 26 than to 48 and its interval widens accordingly. A small negative value is not a paradox: it means questions on one clip disagree slightly more than questions drawn at random, so there is no clustering to correct for.
| Model | Circular accuracy | 95% interval | ICC | pass@1 | 95% interval |
|---|---|---|---|---|---|
| GPT-5.6 Sol | 67% | 52–82% | +0.08 | 75% | 62–87% |
| Kimi K3 | 67% | 50–84% | +0.42 | 73% | 59–85% |
| GPT-5.5 | 65% | 48–82% | +0.42 | 75% | 62–87% |
| Gemini 3.6 Flash | 62% | 46–78% | +0.18 | 70% | 56–84% |
| Gemini 3.7 Flash | 60% | 45–75% | +0.15 | 72% | 59–84% |
| GPT-5.6 Terra | 54% | 37–71% | +0.27 | 71% | 58–84% |
| Claude Opus 5 | 50% | 32–68% | +0.41 | 63% | 48–77% |
| Qwen 3.7 Plus | 46% | 31–61% | −0.02 | 58% | 48–69% |
| Grok 4.3 | 42% | 26–58% | +0.20 | 57% | 43–70% |
| Claude Sonnet 4.6 | 40% | 26–54% | −0.04 | 49% | 36–61% |
| Claude Opus 4.8 | 27% | 14–40% | 0.00 | 41% | 29–53% |
| Gemini 3.1 Pro | 25% | 11–39% | −0.01 | 56% | 46–67% |
The ICC column and the circular intervals come from a cluster-robust variance with the video as the unit; the pass@1 intervals come from a cluster bootstrap, resampling the 26 recordings twenty thousand times, which respects the same unit without assuming a distribution for the fractional scores. The pass@1 intervals run slightly narrower, about ±12 against ±16 near the middle of the table, because a share of orderings is a less noisy reading per question than all-or-nothing. Read the ICC column as a property of the model on this set rather than of the set itself: the same 26 recordings give +0.42 for one model and −0.04 for another.
Consistency per model
Consistency asks a different question from accuracy: not whether the model was right, but whether it gave the same answer at all. Every question is put once per option ordering, four to six times depending on how many options it carries, so each row below is read over the same 48 questions as the board.
| Model | Circular accuracy | pass@1 | Same answer every ordering | Modal answer share | Distinct answers per question |
|---|---|---|---|---|---|
| GPT-5.6 Sol | 67% | 75% | 79% | 93% | 1.23 |
| Kimi K3 | 67% | 73% | 79% | 92% | 1.25 |
| GPT-5.5 | 65% | 75% | 71% | 91% | 1.33 |
| Gemini 3.6 Flash | 62% | 70% | 75% | 92% | 1.33 |
| Gemini 3.7 Flash | 60% | 72% | 69% | 88% | 1.42 |
| GPT-5.6 Terra | 54% | 71% | 60% | 87% | 1.50 |
| Claude Opus 5 | 50% | 63% | 73% | 91% | 1.31 |
| Qwen 3.7 Plus | 46% | 58% | 60% | 86% | 1.48 |
| Grok 4.3 | 42% | 57% | 62% | 84% | 1.50 |
| Claude Sonnet 4.6 | 40% | 49% | 71% | 92% | 1.31 |
| Claude Opus 4.8 | 27% | 41% | 62% | 87% | 1.48 |
| Gemini 3.1 Pro | 25% | 56% | 35% | 78% | 1.83 |
The pass@1 column scores each question as the share of its orderings answered correctly, the unbiased estimate of one attempt: unlike accuracy on the annotator’s original ordering, it favors no particular draw, and the difference is not hypothetical, since GPT-5.5 answers 83% on the original ordering against 75% averaged over all of them. We also checked that the orderings themselves are interchangeable: pooled across the twelve models, the orderings nearly every question carries land within 1.3 points of one another, so no rotation is systematically harder. Everything between a model’s average and its circular score is therefore inconsistency, and that cost runs from 7 points (Kimi K3) to 31 (Gemini 3.1 Pro).
The same answer every ordering column is the strict measure quoted in the text, and it is all-or-nothing: one ordering answered differently and the question does not count, however many agreed. The modal answer share is the same evidence read gently, the average share of a question’s orderings that landed on whichever answer that model gave most often, and the last column says it in plain units. The gentler columns move every model up and change nothing about the order: Gemini 3.1 Pro is last on all three.
Consistency and accuracy rise together at a rank correlation of 0.71, so the rows that break the pattern are where the information is. Claude Opus 5 is steadier than GPT-5.5 and 15 points behind it; Claude Sonnet 4.6 matches GPT-5.5 exactly and is 25 points behind. Both are holding an answer firmly and holding the wrong one, which is a different failure from Gemini 3.1 Pro at the bottom, where the answer itself moves with the option order.
How the human baseline was measured
Five testers worked the set independently, answering the same multiple-choice questions the models were given, from the same recordings, once per question. Nobody was shown a question they had any prior contact with: where a tester had already contributed to one on this project, as its annotator or as its reviewer, that question was withheld from them rather than counted. So each tester saw only footage and questions that were new to them, and the exclusions are the reason two of them answered fewer than 48.
| Tester | Correct | Withheld for prior involvement | Rate |
|---|---|---|---|
| Human 1 | 41 of 48 | 0 | 85% |
| Human 2 | 40 of 47 | 1 | 85% |
| Human 3 | 41 of 48 | 0 | 85% |
| Human 4 | 42 of 48 | 0 | 88% |
| Human 5 | 36 of 43 | 5 | 84% |
| All five | 200 of 234 | 6 | 85% |
Six question-tester pairs were withheld that way, one for the second tester and five for the fifth, leaving 234 answers of a possible 240. Nothing was skipped: every question a tester was eligible for, they answered. Averaging the five rates gives 85.4% and pooling every answer gives 85.5%, so the figure is 85% either way, and the spread across testers is narrow: 1.4 points of standard deviation, from 84% to 88%.
What it can and cannot be compared with. Each tester saw every question once, so 85% is a single-pass score, and it compares to the headline pass@1, where the strongest models reach 75%. It does not compare to circular accuracy, which requires the same question to be answered correctly under every ordering of its options, something no human here was asked to do. Read against circular accuracy the gap is 18 points, and the difference between those two readings is consistency under re-ordering, which this baseline does not measure. Five testers is also a small panel, enough to establish that the questions are answerable and not enough to place the human number precisely.