all leaderboards
vetto · computer anthology

Terminal Tasks v1.0

Every number below comes from one protocol: 5 independent trials per task per configuration, 500 trials per configuration, run on Harbor and graded by each task's own verifier.

100 tasks pass@1 · 5 trials/task 40 model + agent pairs updated September 2026
read the research post
Models
Harness

14 / 40 configurations shown

Use the filter to add or remove configurations.

official board

Pass rate by configuration

pass@1 over 500 trials per configuration (5 per task), with 95% bootstrap confidence intervals. The rank is the published order; the filter hides rows, it does not re-rank them.

# model 0–80%
1 GPT-6 Astra OpenAI · medium · Terminus-2 67.0%
2 GPT-6 Astra OpenAI · high · Codex CLI 65.6%
3 Claude Fable 5.1 Anthropic · high · Terminus-2 62.4%
4 Claude Opus 5 Anthropic · high · Terminus-2 62.0%
5 GPT-5.6 Sol OpenAI · high · Codex CLI 58.0%
6 Claude Fable 5.1 Anthropic · high · Claude Code 55.2%
7 Claude Fable 5 Anthropic · high · Claude Code 54.0%
8 Claude Opus 5 Anthropic · high · Claude Code 54.0%
9 Gemini 3.8 Flash Google DeepMind · medium · Terminus-2 51.8%
10 GPT-5.5 OpenAI · high · Codex CLI 51.6%
11 GPT-5.6 Sol OpenAI · medium · Terminus-2 51.6%
12 Claude Fable 5 Anthropic · high · Terminus-2 50.8%
13 Grok 4.6 xAI · high · Terminus-2 49.2%
14 Gemini 3.7 Flash Google DeepMind · medium · Terminus-2 46.0%
15 Kimi K3 Moonshot AI · max · Terminus-2 40.0%
16 MuseSpark 1.1 Meta · Terminus-2 39.0%
17 GPT-5.5 OpenAI · medium · Terminus-2 38.6%
18 Claude Opus 4.8 Anthropic · high · Claude Code 38.0%
19 GPT-5.4 OpenAI · high · Codex CLI 38.0%
20 Tencent Hy4 Tencent · high · Terminus-2 33.0%
21 DeepSeek V4 Pro (0813) DeepSeek · high · Terminus-2 32.6%
22 GLM-5.2 Z.ai · high · Terminus-2 31.4%
23 Claude Opus 4.8 Anthropic · high · Terminus-2 29.8%
24 Qwen3.8 Flash Alibaba · xhigh · Terminus-2 28.8%
25 Qwen3.8 Max Alibaba · xhigh · Terminus-2 26.0%
26 Claude Sonnet 5 Anthropic · high · Claude Code 25.6%
27 Gemini 3.5 Flash Google DeepMind · high · Gemini CLI 24.4%
28 Gemini 3.5 Flash Google DeepMind · medium · Terminus-2 22.8%
29 GPT-5.4 Mini OpenAI · high · Codex CLI 21.8%
30 Gemini 3.1 Pro Google DeepMind · high · Gemini CLI 19.6%
31 Tencent Hy3 Tencent · high · Terminus-2 18.0%
32 Gemini 3.1 Pro Google DeepMind · high · Terminus-2 17.6%
33 Kimi K2.7 Code Moonshot AI · thinking · Terminus-2 15.4%
34 Claude Opus 4.6 Anthropic · high · Terminus-2 14.8%
35 GPT-5.4 OpenAI · none · Terminus-2 12.0%
36 DeepSeek V4 Pro DeepSeek · high · Terminus-2 10.4%
37 Qwen3.7 Max Alibaba · reasoning · Terminus-2 9.8%
38 Muse Glimmer 30B Meta · high · Terminus-2 5.8%
39 GPT-5.4 Mini OpenAI · none · Terminus-2 4.0%
40 Devstral 2 Mistral · Terminus-2 2.8%

pass@1 over 500 trials per configuration (5 per task), with 95% bootstrap confidence intervals. Ties keep the published order.

breakdown

Pass rate against spend

One dot per configuration, coloured by provider. Switch the x axis between output tokens, cost per trial, median steps and published parameter count; a configuration below 80% telemetry coverage on a metric is named, not plotted.

Pass rate against output tokens
Pass rate against output tokensAll 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair. pass@1 Mean output tokens per trial 0% 10% 20% 30% 40% 50% 60% 70% 1000 10000 100000 1000000 OpenAI Anthropic Google DeepMind xAI Moonshot AI Meta Tencent DeepSeek Z.ai Alibaba Mistral

39 configurations · 500 trials each · x axis log scale. Below 80% telemetry coverage, so not plotted: Gemini 3.1 Pro + Gemini CLI.

Pass rate against cost per trial
Pass rate against cost per trialAll 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair. pass@1 Mean cost per trial 0% 10% 20% 30% 40% 50% 60% 70% 0 USD 0.5 USD 1 USD 1.5 USD 2 USD 2.5 USD 3 USD OpenAI Anthropic Google DeepMind xAI Moonshot AI Meta Tencent DeepSeek Z.ai Alibaba Mistral

39 configurations · 500 trials each · x axis linear. Below 80% telemetry coverage, so not plotted: Gemini 3.1 Pro + Gemini CLI.

Pass rate against median steps
Pass rate against median stepsAll 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair. pass@1 Median steps per trial (harness-defined granularity) 0% 10% 20% 30% 40% 50% 60% 70% 0 20 40 60 80 100 OpenAI Anthropic Google DeepMind xAI Moonshot AI Meta Tencent DeepSeek Z.ai Alibaba Mistral

39 configurations · 500 trials each · x axis linear. Below 80% telemetry coverage, so not plotted: Gemini 3.1 Pro + Gemini CLI.

Pass rate against parameters
Pass rate against parametersThe 10 configurations whose model publishes a parameter count; the closed models have none to plot. pass@1 Total parameters (billions) 0% 10% 20% 30% 40% 10B 100B 1000B 10000B Moonshot AI Tencent Z.ai Alibaba DeepSeek Meta Mistral

10 configurations · 500 trials each · x axis log scale.

Claude Fable — Anthropic's safety classifier can refuse to run Fable; when it does, the trial continues on Anthropic's default fallback model, which may affect its measured performance.

breakdown

Pass rate by category

The share of each configuration's trials that passed the tasks of a category. Categories with fewer than six tasks are left out and named in the caption.

pass@1 by category and configuration
pass@1 by category and configurationRows are the 40 configurations sorted by overall pass@1, columns the 9 categories with at least 6 tasks, hardest first. Hover a cell for the full configuration name. configuration category GPT-6 Astra… GPT-6 Astra… Claude Fabl… Claude Opus… GPT-5.6 Sol… Claude Fabl… Claude Fabl… Claude Opus… Gemini 3.8 … GPT-5.5 + C… GPT-5.6 Sol… Claude Fabl… Grok 4.6 + … Gemini 3.7 … Kimi K3 + T… MuseSpark 1… GPT-5.5 + T… Claude Opus… GPT-5.4 + C… Tencent Hy4… DeepSeek V4… GLM-5.2 + T… Claude Opus… Qwen3.8 Fla… Qwen3.8 Max… Claude Sonn… Gemini 3.5 … Gemini 3.5 … GPT-5.4 Min… Gemini 3.1 … Tencent Hy3… Gemini 3.1 … Kimi K2.7 C… Claude Opus… GPT-5.4 + T… DeepSeek V4… Qwen3.7 Max… Muse Glimme… GPT-5.4 Min… Devstral 2 … Se… De… Ma… Mo… Da… Da… Da… Op… So… 60% 74.3% 51.4% 58.2% 64.4% 83.3% 77.5% 76% 76.7% 51.4% 77.1% 57.1% 49.1% 60% 86.7% 80% 70% 83.3% 45.7% 65.7% 60% 63.6% 66.7% 76.7% 52.5% 70% 56.7% 62.9% 51.4% 55.7% 72.7% 75.6% 56.7% 20% 68% 73.3% 48.6% 48.6% 55.7% 43.6% 40% 86.7% 65% 58% 76.7% 40% 65.7% 47.1% 49.1% 66.7% 40% 52.5% 80% 66.7% 25.7% 20% 48.6% 67.3% 68.9% 63.3% 37.5% 76% 56.7% 42.9% 25.7% 51.4% 49.1% 62.2% 53.3% 35% 74% 60% 17.1% 51.4% 44.3% 34.5% 55.6% 33.3% 77.5% 58% 70% 34.3% 22.9% 55.7% 52.7% 53.3% 50% 60% 54% 66.7% 48.6% 28.6% 40% 58.2% 51.1% 70% 72.5% 54% 26.7% 37.1% 40% 37.1% 50.9% 46.7% 46.7% 42.5% 60% 73.3% 31.4% 57.1% 48.6% 32.7% 40% 53.3% 57.5% 56% 56.7% 17.1% 71.4% 34.3% 30.9% 48.9% 46.7% 52.5% 36% 66.7% 37.1% 42.9% 35.7% 32.7% 37.8% 40% 45% 52% 46.7% 31.4% 37.1% 35.7% 25.5% 37.8% 50% 45% 58% 26.7% 22.9% 17.1% 35.7% 30.9% 33.3% 46.7% 52.5% 30% 30% 34.3% 17.1% 34.3% 36.4% 40% 40% 42.5% 48% 36.7% 34.3% 28.6% 54.3% 25.5% 24.4% 46.7% 57.5% 38% 40% 37.1% 45.7% 42.9% 38.2% 24.4% 10% 32.5% 22% 40% 31.4% 22.9% 14.3% 30.9% 40% 33.3% 52.5% 18% 56.7% 37.1% 0% 18.6% 36.4% 15.6% 36.7% 47.5% 32% 40% 20% 31.4% 17.1% 21.8% 26.7% 40% 22.5% 40% 40% 37.1% 42.9% 24.3% 32.7% 24.4% 30% 30% 20% 33.3% 17.1% 28.6% 30% 18.2% 28.9% 23.3% 42.5% 16% 16.7% 17.1% 8.6% 11.4% 18.2% 33.3% 36.7% 25% 28% 60% 22.9% 5.7% 11.4% 25.5% 37.8% 23.3% 25% 14% 40% 8.6% 11.4% 8.6% 18.2% 37.8% 33.3% 12.5% 14% 20% 22.9% 8.6% 17.1% 9.1% 17.8% 6.7% 42.5% 32% 30% 2.9% 14.3% 11.4% 16.4% 26.7% 23.3% 10% 18% 33.3% 22.9% 14.3% 11.4% 10.9% 13.3% 10% 22.5% 16% 33.3% 5.7% 2.9% 5.7% 14.5% 22.2% 23.3% 10% 10% 33.3% 5.7% 11.4% 8.6% 10.9% 13.3% 16.7% 10% 36% 26.7% 11.4% 0% 11.4% 18.2% 4.4% 13.3% 5% 40% 6.7% 17.1% 2.9% 7.1% 14.5% 13.3% 10% 12.5% 16% 3.3% 5.7% 11.4% 7.1% 23.6% 11.1% 16.7% 5% 8% 3.3% 5.7% 17.1% 11.4% 9.1% 8.9% 6.7% 2.5% 18% 0% 2.9% 0% 7.1% 1.8% 6.7% 0% 27.5% 4% 0% 0% 0% 5.7% 1.8% 6.7% 0% 12.5% 6% 0% 0% 2.9% 0% 0% 2.2% 0% 0% 4% 0% 0% 100%

Task counts per category: Security 7 · Debugging 7 · Machine Learning 14 · Model Training 11 · Data Science 9 · Data Querying 6 · Data Processing 8 · Optimization 10 · Software Engineering 6. Cell = share of that configuration's trials that passed the category's tasks. Left out, under 6 tasks each: File Operations (3), Games (2), Mathematics (2), Personal Assistant (2), Scientific Computing (2), System Administration (5), Video Processing (4), Web Browsing (2) — 22 tasks in all.