all leaderboards
vetto · computer anthology

ProgramBench Vetted

Fifty held-out programs reconstructed from runnable binaries. Every result is pass@1 using the default ProgramBench mini-swe-agent configuration and deterministic behavioral grading.

50 held-out tasks pass@1 17 evaluated models updated August 2026
read the research post
official board

Model results

Three numbers, one ranking: pick the metric and the board re-ranks by it. Resolved is the share of tasks passing every test, almost the share passing at least 95%, and mean reward the macro average of each task’s fractional test score — so a model can grade well on partial progress and still resolve almost nothing.

# model 0–20%
1 Claude Opus 5 (xhigh) Anthropic 14.0%50.0%81.9%
2 Grok 4.6 (xhigh) SpaceXAI 6.0%24.0%68.7%
3 GPT 5.5 (xhigh) OpenAI 2.0%28.0%85.2%
4 DeepSeek V4 Pro 0813 (xhigh) DeepSeek 2.0%22.0%71.8%
5 Claude Sonnet 5 (xhigh) Anthropic 2.0%14.0%57.9%
6 Muse Spark 1.2 (high) Meta 2.0%6.0%55.5%
7 GPT 5.6 Sol (xhigh) OpenAI 0.0%28.0%82.8%
8 Kimi K3 (max) Moonshot AI 0.0%10.0%22.1%
9 Gemini 3.7 Flash (high) Google DeepMind 0.0%6.0%43.4%
10 DeepSeek V4 Flash 0731 (high) DeepSeek 0.0%4.0%51.6%
11 GPT 5.6 Terra (xhigh) OpenAI 0.0%4.0%70.6%
12 GLM-5.2 (xhigh) Z.ai 0.0%4.0%59.1%
13 Gemini 3.6 Flash (high) Google DeepMind 0.0%0.0%57.8%
14 GPT 5.6 Sol (medium) OpenAI 0.0%0.0%61.2%
15 Gemini 3.1 Pro Preview (high) Google DeepMind 0.0%0.0%29.1%
16 GPT 5.6 Terra (medium) OpenAI 0.0%0.0%26.1%
17 Muse Glimmer 30B (high) Meta 0.0%0.0%7.2%

Resolved is the share of the 50 tasks passing 100% of their tests, pass@1 on the default ProgramBench mini-swe-agent configuration. Ties are broken by almost, then by the published order.