Hi — thanks for publishing this. Holding the model, tasks, and runtime fixed while varying only the harness is the comparison that was missing.
A question about the lightest possible path to getting a harness into the results.
The self-serve flow is heavy: you need your own Linux box, a Runta CLI login, a provider API key, a golden-checkpoint provision, a full 30-task trial run, then scoring and reporting — and even then the run defaults to methodology_comparable: false with no leaderboard rank unless you also run a matched control. That is a lot to ask of a harness author who just wants a number next to the published rows.
Is there — or has there been discussion of — a simpler, fully hosted option along these lines?
- The author submits only a harness repo URL and a pinned commit, plus the minimal build/provider settings.
- They pay a fixed fee by scanning a QR code (whatever is easiest: WeChat Pay / Alipay / card).
- An official agent, run by the maintainers with the same skill, scripts, and egress policy as the published baselines, provisions the checkpoint, runs the tasks from an identical fresh restore, scores them, and publishes the harness as a new row under the existing methodology.
That would reduce the user-facing workflow to "here is my repo URL, here is my payment, here is my result," while keeping the run authoritative because the maintainers execute it under the same invariants instead of the author self-reporting.
A few concrete questions:
- Is this already possible, or planned?
- What is the cheapest first step? e.g. accept just the repo URL and run one Terminal-Bench plus one DeepSWE task as a paid smoke test, then the full 30-task set once the plumbing is confirmed.
- How would pricing map to a real run (compute + provider tokens)? Even a rough range would help authors decide.
- If a submitted harness cannot be driven headlessly (no CLI/headless mode, or it needs interactive auth), is that a hard reject or handled case by case?
If self-run is intended to stay the only path, a short README note on why (cost, egress policy, verification) would also help set expectations.
Thanks.
Hi — thanks for publishing this. Holding the model, tasks, and runtime fixed while varying only the harness is the comparison that was missing.
A question about the lightest possible path to getting a harness into the results.
The self-serve flow is heavy: you need your own Linux box, a Runta CLI login, a provider API key, a golden-checkpoint provision, a full 30-task trial run, then scoring and reporting — and even then the run defaults to
methodology_comparable: falsewith no leaderboard rank unless you also run a matched control. That is a lot to ask of a harness author who just wants a number next to the published rows.Is there — or has there been discussion of — a simpler, fully hosted option along these lines?
That would reduce the user-facing workflow to "here is my repo URL, here is my payment, here is my result," while keeping the run authoritative because the maintainers execute it under the same invariants instead of the author self-reporting.
A few concrete questions:
If self-run is intended to stay the only path, a short README note on why (cost, egress policy, verification) would also help set expectations.
Thanks.