Hard and fair benchmarks, scored live.
We build verifier-graded benchmarks with real headroom, and re-score model + agent pairs as new models ship.
Can an agent reconstruct a program from its runnable binary? Fifty held-out tasks measure partial progress and complete resolution with deterministic behavioral tests.
The set spans 17 categories and 11 languages and runtimes, weighted toward data science, ML, optimization, and model-training work, and it grades whether the artifact works within the stated constraints, not code quality or style.
Gullibility, measured: on tiny tasks where a convenient stored field lies and the truth sits one derivation away in the same data, does the model check or take the number it was handed? 50 held-out multi-hop tasks rendered in three substrates, scored pass@1 over five runs.
Video perception, measured: 48 curated multiple-choice questions over 26 recordings filmed to order, each built around one twist. Twelve frontier models scored pass@1 across every ordering of the options, against a five-person human baseline of 85%.
More entries are coming over the next weeks and months, each held to the same standards of deterministic verification, calibrated difficulty, uncertainty-aware reporting, and held-out tasks. Read the Computer Anthology announcement.