Skip to content

### Request: add IntelligenceLab/Long-Horizon-Terminal-Bench to the benchmark allow-list #2648

Description

@zli12321

Hi Hugging Face team,

Following the Evaluation Results docs, I've set up our dataset as a benchmark and would like to request that it be added to the current benchmark allow-list.

Dataset: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench

Benchmark configuration

eval.yaml is in the dataset repo root and validates successfully on the Hub.

  • evaluation_framework: harbor
    • same framework already used by harborframework/terminal-bench-2.0
    • we are not requesting a new framework registration
  • benchmark task:
    • id: lhtb
    • config: tasks
    • split: test

What this benchmark is

LHTB (Long-Horizon-Terminal-Bench) is a 46-task benchmark for evaluating how well LLM agents can sustain useful work in a containerized terminal over hundreds of steps.

Each task is packaged as a Harbor / Terminal-Bench 2.0 task folder, and grading is based on hidden rebuild-from-artifact verifiers rather than self-reported progress.

Relevant links:

Existing eval results

We already have community .eval_results PRs opened against model repos that reference this benchmark with task_id: lhtb, so leaderboard entries are ready once the dataset is allow-listed:

Request

Could you please add IntelligenceLab/Long-Horizon-Terminal-Bench to the benchmark allow-list so that the Benchmark badge and leaderboard can render on the dataset page?

Happy to provide any additional information if helpful. Thanks very much!

cc @julien-c @NathanHB

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions