vetto · computer anthology
ProgramBench Vetted
Fifty held-out programs reconstructed from runnable binaries. Every result is pass@1 using the default ProgramBench mini-swe-agent configuration and deterministic behavioral grading.
official board
Mean reward by model
Ranked by resolved tasks, then by almost-resolved. The bar is mean reward, the macro average of each task’s fractional test score, so a model can grade well on partial progress and still resolve almost nothing.
# model reward almostresolved
1 Claude Opus 5 (xhigh) Anthropic 81.9% 50.0%14.0%
2 Grok 4.6 (xhigh) SpaceXAI 68.7% 24.0%6.0%
3 GPT 5.5 (xhigh) OpenAI 85.2% 28.0%2.0%
4 DeepSeek V4 Pro 0813 (xhigh) DeepSeek 71.8% 22.0%2.0%
5 Claude Sonnet 5 (xhigh) Anthropic 57.9% 14.0%2.0%
6 Muse Spark 1.2 (high) Meta 55.5% 6.0%2.0%
7 GPT 5.6 Sol (xhigh) OpenAI 82.8% 28.0%0.0%
8 Kimi K3 (max) Moonshot AI 22.1% 10.0%0.0%
9 Gemini 3.7 Flash (high) Google DeepMind 43.4% 6.0%0.0%
10 DeepSeek V4 Flash 0731 (high) DeepSeek 51.6% 4.0%0.0%
11 GPT 5.6 Terra (xhigh) OpenAI 70.6% 4.0%0.0%
12 GLM-5.2 (xhigh) Z.ai 59.1% 4.0%0.0%
13 Gemini 3.6 Flash (high) Google DeepMind 57.8% 0.0%0.0%
14 GPT 5.6 Sol (medium) OpenAI 61.2% 0.0%0.0%
15 Gemini 3.1 Pro Preview (high) Google DeepMind 29.1% 0.0%0.0%
16 GPT 5.6 Terra (medium) OpenAI 26.1% 0.0%0.0%
17 Muse Glimmer 30B (high) Meta 7.2% 0.0%0.0%
Sorted by resolved, then almost. Mean reward is the macro average of each task’s fractional test score. Almost is the share of tasks reaching at least 95%; resolved is the share reaching 100%.