all leaderboards
vetto · computer anthology

ProgramBench Vetted

Fifty held-out programs reconstructed from runnable binaries. Every result is pass@1 using the default ProgramBench mini-swe-agent configuration and deterministic behavioral grading.

50 held-out tasks pass@1 17 evaluated models updated August 2026
official board

Mean reward by model

Ranked by resolved tasks, then by almost-resolved. The bar is mean reward, the macro average of each task’s fractional test score, so a model can grade well on partial progress and still resolve almost nothing.

#modelmean reward (0–100%)rewardalmostresolved
1Claude Opus 5xhigh · Anthropic81.9%50.0%14.0%
2Grok 4.6xhigh · SpaceXAI68.7%24.0%6.0%
3GPT 5.5xhigh · OpenAI85.2%28.0%2.0%
4DeepSeek V4 Pro 0813xhigh · DeepSeek71.8%22.0%2.0%
5Claude Sonnet 5xhigh · Anthropic57.9%14.0%2.0%
6Muse Spark 1.2high · Meta55.5%6.0%2.0%
7GPT 5.6 Solxhigh · OpenAI82.8%28.0%0.0%
8Kimi K3max · Moonshot AI22.1%10.0%0.0%
9Gemini 3.7 Flashhigh · Google DeepMind43.4%6.0%0.0%
10DeepSeek V4 Flash 0731high · DeepSeek51.6%4.0%0.0%
11GPT 5.6 Terraxhigh · OpenAI70.6%4.0%0.0%
12GLM-5.2xhigh · Z.ai59.1%4.0%0.0%
13Gemini 3.6 Flashhigh · Google DeepMind57.8%0.0%0.0%
14GPT 5.6 Solmedium · OpenAI61.2%0.0%0.0%
15Gemini 3.1 Pro Previewhigh · Google DeepMind29.1%0.0%0.0%
16GPT 5.6 Terramedium · OpenAI26.1%0.0%0.0%
17Muse Glimmer 30Bhigh · Meta7.2%0.0%0.0%

Sorted by resolved, then almost. Mean reward is the macro average of each task’s fractional test score. Almost is the share of tasks reaching at least 95%; resolved is the share reaching 100%.