← all leaderboards
vetto · computer anthology
Terminal Tasks v1.0
Every number below comes from one protocol: 5 independent trials per task per configuration, 500 trials per configuration, run on Harbor and graded by each task’s own verifier.
Pass@1 is the only metric this board ranks by: each trial is graded all-or-nothing by the task’s own verifier, so there is no almost to report and no fractional reward to average.
Claude Fable — Anthropic’s safety classifier can refuse to run Fable; when it does, the trial continues on Anthropic’s default fallback model, which may affect its measured performance.