Hard and fair benchmarks, scored live.
We build verifier-graded benchmarks with real headroom, and re-score model + agent pairs as new models ship.
Can an agent reconstruct a program from its runnable binary? Fifty held-out tasks measure partial progress and complete resolution with deterministic behavioral tests.
Gullibility, measured: on tiny tasks where a convenient stored field lies and the truth sits one derivation away in the same data, does the model check or take the number it was handed? 50 held-out multi-hop tasks rendered in three substrates, scored pass@1 over five runs.
The set spans 17 categories and 11 languages and runtimes, weighted toward data science, ML, optimization, and model-training work, and it grades whether the artifact works within the stated constraints, not code quality or style.
More entries are coming over the next weeks and months, each held to the same standards of deterministic verification, calibrated difficulty, uncertainty-aware reporting, and held-out tasks. Read the Computer Anthology announcement.