vetto · computer anthology
Terminal Tasks v1.0 Every number below comes from one protocol: 5 independent trials per task per configuration, 500 trials per configuration, run on Harbor and graded by each task's own verifier.
100 tasks pass@1 · 5 trials/task 40 model + agent pairs updated September 2026
Models GPT-6 Astra, Claude Fable 5.1 +9 more
14 / 40 configurations shown
Use the filter to add or remove configurations.
official board
Pass rate by configuration pass@1 over 500 trials per configuration (5 per task), with 95% bootstrap confidence intervals. The rank is the published order; the filter hides rows, it does not re-rank them.
pass@1 over 500 trials per configuration (5 per task), with 95% bootstrap confidence intervals. Ties keep the published order.
breakdown
Pass rate against spend One dot per configuration, coloured by provider. Switch the x axis between output tokens, cost per trial, median steps and published parameter count; a configuration below 80% telemetry coverage on a metric is named, not plotted.
Output tokens Cost per trial Median steps Parameters
Pass rate against output tokens Pass rate against output tokens All 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair.
pass@1
Mean output tokens per trial
0%
10%
20%
30%
40%
50%
60%
70%
1000
10000
100000
1000000
OpenAI
Anthropic
Google DeepMind
xAI
Moonshot AI
Meta
Tencent
DeepSeek
Z.ai
Alibaba
Mistral
Pass rate against output tokens All 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair.
pass@1
Mean output tokens per trial
0%
10%
20%
30%
40%
50%
60%
70%
1000
10000
100000
1000000
OpenAI
Anthropic
Google DeepMind
xAI
Moonshot AI
Meta
Tencent
DeepSeek
Z.ai
Alibaba
Mistral
Pass rate against output tokens All 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair.
pass@1
Mean output tokens per trial
0%
10%
20%
30%
40%
50%
60%
70%
1000
10000
100000
1000000
OpenAI
Anthropic
Google DeepMind
xAI
Moonshot AI
Meta
Tencent
DeepSeek
Z.ai
Alibaba
Mistral
Pass rate against output tokens All 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair.
pass@1
Mean output tokens per trial
0%
10%
20%
30%
40%
50%
60%
70%
1000
10000
100000
1000000
OpenAI
Anthropic
Google DeepMind
xAI
Moonshot AI
Meta
Tencent
DeepSeek
Z.ai
Alibaba
Mistral
39 configurations · 500 trials each · x axis log scale. Below 80% telemetry coverage, so not plotted: Gemini 3.1 Pro + Gemini CLI.
Pass rate against cost per trial Pass rate against cost per trial All 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair.
pass@1
Mean cost per trial
0%
10%
20%
30%
40%
50%
60%
70%
0 USD
0.5 USD
1 USD
1.5 USD
2 USD
2.5 USD
3 USD
OpenAI
Anthropic
Google DeepMind
xAI
Moonshot AI
Meta
Tencent
DeepSeek
Z.ai
Alibaba
Mistral
Pass rate against cost per trial All 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair.
pass@1
Mean cost per trial
0%
10%
20%
30%
40%
50%
60%
70%
0 USD
0.5 USD
1 USD
1.5 USD
2 USD
2.5 USD
3 USD
OpenAI
Anthropic
Google DeepMind
xAI
Moonshot AI
Meta
Tencent
DeepSeek
Z.ai
Alibaba
Mistral
Pass rate against cost per trial All 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair.
pass@1
Mean cost per trial
0%
10%
20%
30%
40%
50%
60%
70%
0 USD
0.5 USD
1 USD
1.5 USD
2 USD
2.5 USD
3 USD
OpenAI
Anthropic
Google DeepMind
xAI
Moonshot AI
Meta
Tencent
DeepSeek
Z.ai
Alibaba
Mistral
Pass rate against cost per trial All 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair.
pass@1
Mean cost per trial
0%
10%
20%
30%
40%
50%
60%
70%
0 USD
0.5 USD
1 USD
1.5 USD
2 USD
2.5 USD
3 USD
OpenAI
Anthropic
Google DeepMind
xAI
Moonshot AI
Meta
Tencent
DeepSeek
Z.ai
Alibaba
Mistral
39 configurations · 500 trials each · x axis linear. Below 80% telemetry coverage, so not plotted: Gemini 3.1 Pro + Gemini CLI.
Pass rate against median steps Pass rate against median steps All 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair.
pass@1
Median steps per trial (harness-defined granularity)
0%
10%
20%
30%
40%
50%
60%
70%
0
20
40
60
80
100
OpenAI
Anthropic
Google DeepMind
xAI
Moonshot AI
Meta
Tencent
DeepSeek
Z.ai
Alibaba
Mistral
Pass rate against median steps All 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair.
pass@1
Median steps per trial (harness-defined granularity)
0%
10%
20%
30%
40%
50%
60%
70%
0
20
40
60
80
100
OpenAI
Anthropic
Google DeepMind
xAI
Moonshot AI
Meta
Tencent
DeepSeek
Z.ai
Alibaba
Mistral
Pass rate against median steps All 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair.
pass@1
Median steps per trial (harness-defined granularity)
0%
10%
20%
30%
40%
50%
60%
70%
0
20
40
60
80
100
OpenAI
Anthropic
Google DeepMind
xAI
Moonshot AI
Meta
Tencent
DeepSeek
Z.ai
Alibaba
Mistral
Pass rate against median steps All 40 configurations. One dot per model and harness pair, coloured by provider; hover a dot for the pair.
pass@1
Median steps per trial (harness-defined granularity)
0%
10%
20%
30%
40%
50%
60%
70%
0
20
40
60
80
100
OpenAI
Anthropic
Google DeepMind
xAI
Moonshot AI
Meta
Tencent
DeepSeek
Z.ai
Alibaba
Mistral
39 configurations · 500 trials each · x axis linear. Below 80% telemetry coverage, so not plotted: Gemini 3.1 Pro + Gemini CLI.
Pass rate against parameters Pass rate against parameters The 10 configurations whose model publishes a parameter count; the closed models have none to plot.
pass@1
Total parameters (billions)
0%
10%
20%
30%
40%
10B
100B
1000B
10000B
Moonshot AI
Tencent
Z.ai
Alibaba
DeepSeek
Meta
Mistral
Pass rate against parameters The 10 configurations whose model publishes a parameter count; the closed models have none to plot.
pass@1
Total parameters (billions)
0%
10%
20%
30%
40%
10B
100B
1000B
10000B
Moonshot AI
Tencent
Z.ai
Alibaba
DeepSeek
Meta
Mistral
Pass rate against parameters The 10 configurations whose model publishes a parameter count; the closed models have none to plot.
pass@1
Total parameters (billions)
0%
10%
20%
30%
40%
10B
100B
1000B
10000B
Moonshot AI
Tencent
Z.ai
Alibaba
DeepSeek
Meta
Mistral
Pass rate against parameters The 10 configurations whose model publishes a parameter count; the closed models have none to plot.
pass@1
Total parameters (billions)
0%
10%
20%
30%
40%
10B
100B
1000B
10000B
Moonshot AI
Tencent
Z.ai
Alibaba
DeepSeek
Meta
Mistral
10 configurations · 500 trials each · x axis log scale.
Claude Fable — Anthropic's safety
classifier can refuse to run Fable; when it does, the trial continues on Anthropic's default
fallback model, which may affect its measured performance.
breakdown
Pass rate by category The share of each configuration's trials that passed the tasks of a category. Categories with fewer than six tasks are left out and named in the caption.
pass@1 by category and configuration pass@1 by category and configuration Rows are the 40 configurations sorted by overall pass@1, columns the 9 categories with at least 6 tasks, hardest first. Hover a cell for the full configuration name.
configuration
category
GPT-6 Astra…
GPT-6 Astra…
Claude Fabl…
Claude Opus…
GPT-5.6 Sol…
Claude Fabl…
Claude Fabl…
Claude Opus…
Gemini 3.8 …
GPT-5.5 + C…
GPT-5.6 Sol…
Claude Fabl…
Grok 4.6 + …
Gemini 3.7 …
Kimi K3 + T…
MuseSpark 1…
GPT-5.5 + T…
Claude Opus…
GPT-5.4 + C…
Tencent Hy4…
DeepSeek V4…
GLM-5.2 + T…
Claude Opus…
Qwen3.8 Fla…
Qwen3.8 Max…
Claude Sonn…
Gemini 3.5 …
Gemini 3.5 …
GPT-5.4 Min…
Gemini 3.1 …
Tencent Hy3…
Gemini 3.1 …
Kimi K2.7 C…
Claude Opus…
GPT-5.4 + T…
DeepSeek V4…
Qwen3.7 Max…
Muse Glimme…
GPT-5.4 Min…
Devstral 2 …
Se…
De…
Ma…
Mo…
Da…
Da…
Da…
Op…
So…
60%
74.3%
51.4%
58.2%
64.4%
83.3%
77.5%
76%
76.7%
51.4%
77.1%
57.1%
49.1%
60%
86.7%
80%
70%
83.3%
45.7%
65.7%
60%
63.6%
66.7%
76.7%
52.5%
70%
56.7%
62.9%
51.4%
55.7%
72.7%
75.6%
56.7%
20%
68%
73.3%
48.6%
48.6%
55.7%
43.6%
40%
86.7%
65%
58%
76.7%
40%
65.7%
47.1%
49.1%
66.7%
40%
52.5%
80%
66.7%
25.7%
20%
48.6%
67.3%
68.9%
63.3%
37.5%
76%
56.7%
42.9%
25.7%
51.4%
49.1%
62.2%
53.3%
35%
74%
60%
17.1%
51.4%
44.3%
34.5%
55.6%
33.3%
77.5%
58%
70%
34.3%
22.9%
55.7%
52.7%
53.3%
50%
60%
54%
66.7%
48.6%
28.6%
40%
58.2%
51.1%
70%
72.5%
54%
26.7%
37.1%
40%
37.1%
50.9%
46.7%
46.7%
42.5%
60%
73.3%
31.4%
57.1%
48.6%
32.7%
40%
53.3%
57.5%
56%
56.7%
17.1%
71.4%
34.3%
30.9%
48.9%
46.7%
52.5%
36%
66.7%
37.1%
42.9%
35.7%
32.7%
37.8%
40%
45%
52%
46.7%
31.4%
37.1%
35.7%
25.5%
37.8%
50%
45%
58%
26.7%
22.9%
17.1%
35.7%
30.9%
33.3%
46.7%
52.5%
30%
30%
34.3%
17.1%
34.3%
36.4%
40%
40%
42.5%
48%
36.7%
34.3%
28.6%
54.3%
25.5%
24.4%
46.7%
57.5%
38%
40%
37.1%
45.7%
42.9%
38.2%
24.4%
10%
32.5%
22%
40%
31.4%
22.9%
14.3%
30.9%
40%
33.3%
52.5%
18%
56.7%
37.1%
0%
18.6%
36.4%
15.6%
36.7%
47.5%
32%
40%
20%
31.4%
17.1%
21.8%
26.7%
40%
22.5%
40%
40%
37.1%
42.9%
24.3%
32.7%
24.4%
30%
30%
20%
33.3%
17.1%
28.6%
30%
18.2%
28.9%
23.3%
42.5%
16%
16.7%
17.1%
8.6%
11.4%
18.2%
33.3%
36.7%
25%
28%
60%
22.9%
5.7%
11.4%
25.5%
37.8%
23.3%
25%
14%
40%
8.6%
11.4%
8.6%
18.2%
37.8%
33.3%
12.5%
14%
20%
22.9%
8.6%
17.1%
9.1%
17.8%
6.7%
42.5%
32%
30%
2.9%
14.3%
11.4%
16.4%
26.7%
23.3%
10%
18%
33.3%
22.9%
14.3%
11.4%
10.9%
13.3%
10%
22.5%
16%
33.3%
5.7%
2.9%
5.7%
14.5%
22.2%
23.3%
10%
10%
33.3%
5.7%
11.4%
8.6%
10.9%
13.3%
16.7%
10%
36%
26.7%
11.4%
0%
11.4%
18.2%
4.4%
13.3%
5%
40%
6.7%
17.1%
2.9%
7.1%
14.5%
13.3%
10%
12.5%
16%
3.3%
5.7%
11.4%
7.1%
23.6%
11.1%
16.7%
5%
8%
3.3%
5.7%
17.1%
11.4%
9.1%
8.9%
6.7%
2.5%
18%
0%
2.9%
0%
7.1%
1.8%
6.7%
0%
27.5%
4%
0%
0%
0%
5.7%
1.8%
6.7%
0%
12.5%
6%
0%
0%
2.9%
0%
0%
2.2%
0%
0%
4%
0%
0%
100%
pass@1 by category and configuration Rows are the 40 configurations sorted by overall pass@1, columns the 9 categories with at least 6 tasks, hardest first. Hover a cell for the full configuration name.
configuration
category
GPT-6 Astra…
GPT-6 Astra…
Claude Fabl…
Claude Opus…
GPT-5.6 Sol…
Claude Fabl…
Claude Fabl…
Claude Opus…
Gemini 3.8 …
GPT-5.5 + C…
GPT-5.6 Sol…
Claude Fabl…
Grok 4.6 + …
Gemini 3.7 …
Kimi K3 + T…
MuseSpark 1…
GPT-5.5 + T…
Claude Opus…
GPT-5.4 + C…
Tencent Hy4…
DeepSeek V4…
GLM-5.2 + T…
Claude Opus…
Qwen3.8 Fla…
Qwen3.8 Max…
Claude Sonn…
Gemini 3.5 …
Gemini 3.5 …
GPT-5.4 Min…
Gemini 3.1 …
Tencent Hy3…
Gemini 3.1 …
Kimi K2.7 C…
Claude Opus…
GPT-5.4 + T…
DeepSeek V4…
Qwen3.7 Max…
Muse Glimme…
GPT-5.4 Min…
Devstral 2 …
Se…
De…
Ma…
Mo…
Da…
Da…
Da…
Op…
So…
60%
74.3%
51.4%
58.2%
64.4%
83.3%
77.5%
76%
76.7%
51.4%
77.1%
57.1%
49.1%
60%
86.7%
80%
70%
83.3%
45.7%
65.7%
60%
63.6%
66.7%
76.7%
52.5%
70%
56.7%
62.9%
51.4%
55.7%
72.7%
75.6%
56.7%
20%
68%
73.3%
48.6%
48.6%
55.7%
43.6%
40%
86.7%
65%
58%
76.7%
40%
65.7%
47.1%
49.1%
66.7%
40%
52.5%
80%
66.7%
25.7%
20%
48.6%
67.3%
68.9%
63.3%
37.5%
76%
56.7%
42.9%
25.7%
51.4%
49.1%
62.2%
53.3%
35%
74%
60%
17.1%
51.4%
44.3%
34.5%
55.6%
33.3%
77.5%
58%
70%
34.3%
22.9%
55.7%
52.7%
53.3%
50%
60%
54%
66.7%
48.6%
28.6%
40%
58.2%
51.1%
70%
72.5%
54%
26.7%
37.1%
40%
37.1%
50.9%
46.7%
46.7%
42.5%
60%
73.3%
31.4%
57.1%
48.6%
32.7%
40%
53.3%
57.5%
56%
56.7%
17.1%
71.4%
34.3%
30.9%
48.9%
46.7%
52.5%
36%
66.7%
37.1%
42.9%
35.7%
32.7%
37.8%
40%
45%
52%
46.7%
31.4%
37.1%
35.7%
25.5%
37.8%
50%
45%
58%
26.7%
22.9%
17.1%
35.7%
30.9%
33.3%
46.7%
52.5%
30%
30%
34.3%
17.1%
34.3%
36.4%
40%
40%
42.5%
48%
36.7%
34.3%
28.6%
54.3%
25.5%
24.4%
46.7%
57.5%
38%
40%
37.1%
45.7%
42.9%
38.2%
24.4%
10%
32.5%
22%
40%
31.4%
22.9%
14.3%
30.9%
40%
33.3%
52.5%
18%
56.7%
37.1%
0%
18.6%
36.4%
15.6%
36.7%
47.5%
32%
40%
20%
31.4%
17.1%
21.8%
26.7%
40%
22.5%
40%
40%
37.1%
42.9%
24.3%
32.7%
24.4%
30%
30%
20%
33.3%
17.1%
28.6%
30%
18.2%
28.9%
23.3%
42.5%
16%
16.7%
17.1%
8.6%
11.4%
18.2%
33.3%
36.7%
25%
28%
60%
22.9%
5.7%
11.4%
25.5%
37.8%
23.3%
25%
14%
40%
8.6%
11.4%
8.6%
18.2%
37.8%
33.3%
12.5%
14%
20%
22.9%
8.6%
17.1%
9.1%
17.8%
6.7%
42.5%
32%
30%
2.9%
14.3%
11.4%
16.4%
26.7%
23.3%
10%
18%
33.3%
22.9%
14.3%
11.4%
10.9%
13.3%
10%
22.5%
16%
33.3%
5.7%
2.9%
5.7%
14.5%
22.2%
23.3%
10%
10%
33.3%
5.7%
11.4%
8.6%
10.9%
13.3%
16.7%
10%
36%
26.7%
11.4%
0%
11.4%
18.2%
4.4%
13.3%
5%
40%
6.7%
17.1%
2.9%
7.1%
14.5%
13.3%
10%
12.5%
16%
3.3%
5.7%
11.4%
7.1%
23.6%
11.1%
16.7%
5%
8%
3.3%
5.7%
17.1%
11.4%
9.1%
8.9%
6.7%
2.5%
18%
0%
2.9%
0%
7.1%
1.8%
6.7%
0%
27.5%
4%
0%
0%
0%
5.7%
1.8%
6.7%
0%
12.5%
6%
0%
0%
2.9%
0%
0%
2.2%
0%
0%
4%
0%
0%
100%
pass@1 by category and configuration Rows are the 40 configurations sorted by overall pass@1, columns the 9 categories with at least 6 tasks, hardest first. Hover a cell for the full configuration name.
configuration
category
GPT-6 Astra + Terminu…
GPT-6 Astra + Codex C…
Claude Fable 5.1 + Te…
Claude Opus 5 + Termi…
GPT-5.6 Sol + Codex C…
Claude Fable 5.1 + Cl…
Claude Fable 5 + Clau…
Claude Opus 5 + Claud…
Gemini 3.8 Flash + Te…
GPT-5.5 + Codex CLI
GPT-5.6 Sol + Terminu…
Claude Fable 5 + Term…
Grok 4.6 + Terminus-2
Gemini 3.7 Flash + Te…
Kimi K3 + Terminus-2
MuseSpark 1.1 + Termi…
GPT-5.5 + Terminus-2
Claude Opus 4.8 + Cla…
GPT-5.4 + Codex CLI
Tencent Hy4 + Terminu…
DeepSeek V4 Pro (0813…
GLM-5.2 + Terminus-2
Claude Opus 4.8 + Ter…
Qwen3.8 Flash + Termi…
Qwen3.8 Max + Terminu…
Claude Sonnet 5 + Cla…
Gemini 3.5 Flash + Ge…
Gemini 3.5 Flash + Te…
GPT-5.4 Mini + Codex …
Gemini 3.1 Pro + Gemi…
Tencent Hy3 + Terminu…
Gemini 3.1 Pro + Term…
Kimi K2.7 Code + Term…
Claude Opus 4.6 + Ter…
GPT-5.4 + Terminus-2
DeepSeek V4 Pro + Ter…
Qwen3.7 Max + Terminu…
Muse Glimmer 30B + Te…
GPT-5.4 Mini + Termin…
Devstral 2 + Terminus…
Securi…
Debugg…
Machin…
Model …
Data S…
Data Q…
Data P…
Optimi…
Softwa…
60%
74.3%
51.4%
58.2%
64.4%
83.3%
77.5%
76%
76.7%
51.4%
77.1%
57.1%
49.1%
60%
86.7%
80%
70%
83.3%
45.7%
65.7%
60%
63.6%
66.7%
76.7%
52.5%
70%
56.7%
62.9%
51.4%
55.7%
72.7%
75.6%
56.7%
20%
68%
73.3%
48.6%
48.6%
55.7%
43.6%
40%
86.7%
65%
58%
76.7%
40%
65.7%
47.1%
49.1%
66.7%
40%
52.5%
80%
66.7%
25.7%
20%
48.6%
67.3%
68.9%
63.3%
37.5%
76%
56.7%
42.9%
25.7%
51.4%
49.1%
62.2%
53.3%
35%
74%
60%
17.1%
51.4%
44.3%
34.5%
55.6%
33.3%
77.5%
58%
70%
34.3%
22.9%
55.7%
52.7%
53.3%
50%
60%
54%
66.7%
48.6%
28.6%
40%
58.2%
51.1%
70%
72.5%
54%
26.7%
37.1%
40%
37.1%
50.9%
46.7%
46.7%
42.5%
60%
73.3%
31.4%
57.1%
48.6%
32.7%
40%
53.3%
57.5%
56%
56.7%
17.1%
71.4%
34.3%
30.9%
48.9%
46.7%
52.5%
36%
66.7%
37.1%
42.9%
35.7%
32.7%
37.8%
40%
45%
52%
46.7%
31.4%
37.1%
35.7%
25.5%
37.8%
50%
45%
58%
26.7%
22.9%
17.1%
35.7%
30.9%
33.3%
46.7%
52.5%
30%
30%
34.3%
17.1%
34.3%
36.4%
40%
40%
42.5%
48%
36.7%
34.3%
28.6%
54.3%
25.5%
24.4%
46.7%
57.5%
38%
40%
37.1%
45.7%
42.9%
38.2%
24.4%
10%
32.5%
22%
40%
31.4%
22.9%
14.3%
30.9%
40%
33.3%
52.5%
18%
56.7%
37.1%
0%
18.6%
36.4%
15.6%
36.7%
47.5%
32%
40%
20%
31.4%
17.1%
21.8%
26.7%
40%
22.5%
40%
40%
37.1%
42.9%
24.3%
32.7%
24.4%
30%
30%
20%
33.3%
17.1%
28.6%
30%
18.2%
28.9%
23.3%
42.5%
16%
16.7%
17.1%
8.6%
11.4%
18.2%
33.3%
36.7%
25%
28%
60%
22.9%
5.7%
11.4%
25.5%
37.8%
23.3%
25%
14%
40%
8.6%
11.4%
8.6%
18.2%
37.8%
33.3%
12.5%
14%
20%
22.9%
8.6%
17.1%
9.1%
17.8%
6.7%
42.5%
32%
30%
2.9%
14.3%
11.4%
16.4%
26.7%
23.3%
10%
18%
33.3%
22.9%
14.3%
11.4%
10.9%
13.3%
10%
22.5%
16%
33.3%
5.7%
2.9%
5.7%
14.5%
22.2%
23.3%
10%
10%
33.3%
5.7%
11.4%
8.6%
10.9%
13.3%
16.7%
10%
36%
26.7%
11.4%
0%
11.4%
18.2%
4.4%
13.3%
5%
40%
6.7%
17.1%
2.9%
7.1%
14.5%
13.3%
10%
12.5%
16%
3.3%
5.7%
11.4%
7.1%
23.6%
11.1%
16.7%
5%
8%
3.3%
5.7%
17.1%
11.4%
9.1%
8.9%
6.7%
2.5%
18%
0%
2.9%
0%
7.1%
1.8%
6.7%
0%
27.5%
4%
0%
0%
0%
5.7%
1.8%
6.7%
0%
12.5%
6%
0%
0%
2.9%
0%
0%
2.2%
0%
0%
4%
0%
0%
100%
pass@1 by category and configuration Rows are the 40 configurations sorted by overall pass@1, columns the 9 categories with at least 6 tasks, hardest first. Hover a cell for the full configuration name.
configuration
category
GPT-6 Astra + Terminu…
GPT-6 Astra + Codex C…
Claude Fable 5.1 + Te…
Claude Opus 5 + Termi…
GPT-5.6 Sol + Codex C…
Claude Fable 5.1 + Cl…
Claude Fable 5 + Clau…
Claude Opus 5 + Claud…
Gemini 3.8 Flash + Te…
GPT-5.5 + Codex CLI
GPT-5.6 Sol + Terminu…
Claude Fable 5 + Term…
Grok 4.6 + Terminus-2
Gemini 3.7 Flash + Te…
Kimi K3 + Terminus-2
MuseSpark 1.1 + Termi…
GPT-5.5 + Terminus-2
Claude Opus 4.8 + Cla…
GPT-5.4 + Codex CLI
Tencent Hy4 + Terminu…
DeepSeek V4 Pro (0813…
GLM-5.2 + Terminus-2
Claude Opus 4.8 + Ter…
Qwen3.8 Flash + Termi…
Qwen3.8 Max + Terminu…
Claude Sonnet 5 + Cla…
Gemini 3.5 Flash + Ge…
Gemini 3.5 Flash + Te…
GPT-5.4 Mini + Codex …
Gemini 3.1 Pro + Gemi…
Tencent Hy3 + Terminu…
Gemini 3.1 Pro + Term…
Kimi K2.7 Code + Term…
Claude Opus 4.6 + Ter…
GPT-5.4 + Terminus-2
DeepSeek V4 Pro + Ter…
Qwen3.7 Max + Terminu…
Muse Glimmer 30B + Te…
GPT-5.4 Mini + Termin…
Devstral 2 + Terminus…
Securi…
Debugg…
Machin…
Model …
Data S…
Data Q…
Data P…
Optimi…
Softwa…
60%
74.3%
51.4%
58.2%
64.4%
83.3%
77.5%
76%
76.7%
51.4%
77.1%
57.1%
49.1%
60%
86.7%
80%
70%
83.3%
45.7%
65.7%
60%
63.6%
66.7%
76.7%
52.5%
70%
56.7%
62.9%
51.4%
55.7%
72.7%
75.6%
56.7%
20%
68%
73.3%
48.6%
48.6%
55.7%
43.6%
40%
86.7%
65%
58%
76.7%
40%
65.7%
47.1%
49.1%
66.7%
40%
52.5%
80%
66.7%
25.7%
20%
48.6%
67.3%
68.9%
63.3%
37.5%
76%
56.7%
42.9%
25.7%
51.4%
49.1%
62.2%
53.3%
35%
74%
60%
17.1%
51.4%
44.3%
34.5%
55.6%
33.3%
77.5%
58%
70%
34.3%
22.9%
55.7%
52.7%
53.3%
50%
60%
54%
66.7%
48.6%
28.6%
40%
58.2%
51.1%
70%
72.5%
54%
26.7%
37.1%
40%
37.1%
50.9%
46.7%
46.7%
42.5%
60%
73.3%
31.4%
57.1%
48.6%
32.7%
40%
53.3%
57.5%
56%
56.7%
17.1%
71.4%
34.3%
30.9%
48.9%
46.7%
52.5%
36%
66.7%
37.1%
42.9%
35.7%
32.7%
37.8%
40%
45%
52%
46.7%
31.4%
37.1%
35.7%
25.5%
37.8%
50%
45%
58%
26.7%
22.9%
17.1%
35.7%
30.9%
33.3%
46.7%
52.5%
30%
30%
34.3%
17.1%
34.3%
36.4%
40%
40%
42.5%
48%
36.7%
34.3%
28.6%
54.3%
25.5%
24.4%
46.7%
57.5%
38%
40%
37.1%
45.7%
42.9%
38.2%
24.4%
10%
32.5%
22%
40%
31.4%
22.9%
14.3%
30.9%
40%
33.3%
52.5%
18%
56.7%
37.1%
0%
18.6%
36.4%
15.6%
36.7%
47.5%
32%
40%
20%
31.4%
17.1%
21.8%
26.7%
40%
22.5%
40%
40%
37.1%
42.9%
24.3%
32.7%
24.4%
30%
30%
20%
33.3%
17.1%
28.6%
30%
18.2%
28.9%
23.3%
42.5%
16%
16.7%
17.1%
8.6%
11.4%
18.2%
33.3%
36.7%
25%
28%
60%
22.9%
5.7%
11.4%
25.5%
37.8%
23.3%
25%
14%
40%
8.6%
11.4%
8.6%
18.2%
37.8%
33.3%
12.5%
14%
20%
22.9%
8.6%
17.1%
9.1%
17.8%
6.7%
42.5%
32%
30%
2.9%
14.3%
11.4%
16.4%
26.7%
23.3%
10%
18%
33.3%
22.9%
14.3%
11.4%
10.9%
13.3%
10%
22.5%
16%
33.3%
5.7%
2.9%
5.7%
14.5%
22.2%
23.3%
10%
10%
33.3%
5.7%
11.4%
8.6%
10.9%
13.3%
16.7%
10%
36%
26.7%
11.4%
0%
11.4%
18.2%
4.4%
13.3%
5%
40%
6.7%
17.1%
2.9%
7.1%
14.5%
13.3%
10%
12.5%
16%
3.3%
5.7%
11.4%
7.1%
23.6%
11.1%
16.7%
5%
8%
3.3%
5.7%
17.1%
11.4%
9.1%
8.9%
6.7%
2.5%
18%
0%
2.9%
0%
7.1%
1.8%
6.7%
0%
27.5%
4%
0%
0%
0%
5.7%
1.8%
6.7%
0%
12.5%
6%
0%
0%
2.9%
0%
0%
2.2%
0%
0%
4%
0%
0%
100%
Task counts per category: Security 7 · Debugging 7 · Machine Learning 14 · Model Training 11 · Data Science 9 · Data Querying 6 · Data Processing 8 · Optimization 10 · Software Engineering 6. Cell = share of that configuration's trials that passed the category's tasks. Left out, under 6 tasks each: File Operations (3), Games (2), Mathematics (2), Personal Assistant (2), Scientific Computing (2), System Administration (5), Video Processing (4), Web Browsing (2) — 22 tasks in all.