Leaderboards

Hard and fair benchmarks, scored live.

We build verifier-graded benchmarks with real headroom, and re-score model + agent pairs as new models ship.

Computer Anthology
ProgramBench Vetted

Can an agent reconstruct a program from its runnable binary? Fifty held-out tasks measure partial progress and complete resolution with deterministic behavioral tests.

Top modelresolved
Claude Opus 5 50.0% almost 14.0%
Grok 4.6 24.0% almost 6.0%
GPT 5.5 28.0% almost 2.0%
DeepSeek V4 Pro 0813 22.0% almost 2.0%
Claude Sonnet 5 14.0% almost 2.0%
all 17 models
open the leaderboard
Vetto Research
GulliBench

Gullibility, measured: on tiny tasks where a convenient stored field lies and the truth sits one derivation away in the same data, does the model check or take the number it was handed? 50 held-out multi-hop tasks rendered in three substrates, scored pass@1 over five runs.

Top modelpass@1
Opus 5 95% CI 39–60 49.3%
Fable 5 95% CI 41–55 48.1%
Muse Spark 1.2 95% CI 33–51 42.0%
Gemini 3.1 Pro 95% CI 14–32 22.5%
Kimi K3 95% CI 12–24 17.6%
all 15 models
open the leaderboard
Computer Anthology
Terminal Tasks v1.0

The set spans 17 categories and 11 languages and runtimes, weighted toward data science, ML, optimization, and model-training work, and it grades whether the artifact works within the stated constraints, not code quality or style.

Flagship modelspass@1
Claude Opus 5 Terminus-2 · Anthropic 62.0%
GPT-5.6 Sol Codex CLI · OpenAI 58.0%
Grok 4.6 Terminus-2 · SpaceXAI 49.2%
Gemini 3.7 Flash Terminus-2 · Google DeepMind 46.0%
MuseSpark 1.1 Terminus-2 · Meta 39.0%
all 28 model + agent pairs
open the leaderboard

More entries are coming over the next weeks and months, each held to the same standards of deterministic verification, calibrated difficulty, uncertainty-aware reporting, and held-out tasks. Read the Computer Anthology announcement.