Hi Hugging Face team,
Following the Evaluation Results docs, I've set up our dataset as a benchmark and would like to request that it be added to the current benchmark allow-list.
Dataset: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench
Benchmark configuration
eval.yaml is in the dataset repo root and validates successfully on the Hub.
evaluation_framework: harbor
- same framework already used by
harborframework/terminal-bench-2.0
- we are not requesting a new framework registration
- benchmark task:
id: lhtb
config: tasks
split: test
What this benchmark is
LHTB (Long-Horizon-Terminal-Bench) is a 46-task benchmark for evaluating how well LLM agents can sustain useful work in a containerized terminal over hundreds of steps.
Each task is packaged as a Harbor / Terminal-Bench 2.0 task folder, and grading is based on hidden rebuild-from-artifact verifiers rather than self-reported progress.
Relevant links:
Existing eval results
We already have community .eval_results PRs opened against model repos that reference this benchmark with task_id: lhtb, so leaderboard entries are ready once the dataset is allow-listed:
Request
Could you please add IntelligenceLab/Long-Horizon-Terminal-Bench to the benchmark allow-list so that the Benchmark badge and leaderboard can render on the dataset page?
Happy to provide any additional information if helpful. Thanks very much!
cc @julien-c @NathanHB
Hi Hugging Face team,
Following the Evaluation Results docs, I've set up our dataset as a benchmark and would like to request that it be added to the current benchmark allow-list.
Dataset: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench
Benchmark configuration
eval.yamlis in the dataset repo root and validates successfully on the Hub.evaluation_framework: harborharborframework/terminal-bench-2.0id: lhtbconfig: taskssplit: testWhat this benchmark is
LHTB (Long-Horizon-Terminal-Bench) is a 46-task benchmark for evaluating how well LLM agents can sustain useful work in a containerized terminal over hundreds of steps.
Each task is packaged as a Harbor / Terminal-Bench 2.0 task folder, and grading is based on hidden rebuild-from-artifact verifiers rather than self-reported progress.
Relevant links:
Existing eval results
We already have community
.eval_resultsPRs opened against model repos that reference this benchmark withtask_id: lhtb, so leaderboard entries are ready once the dataset is allow-listed:MiniMaxAI/MiniMax-M3discussion Improve right-to-left language support in inputs and outputs #24deepseek-ai/DeepSeek-V4-Prodiscussion fixes a typoffmeg-->ffmpeg#210moonshotai/Kimi-K2.6discussion Different background color for markdown inline <code> blocks #55moonshotai/Kimi-K2.7-Codediscussion Widget for Table to text generation #13tencent/Hy3discussion Widget for zero-shot image classification #19zai-org/GLM-5.1discussion Tracking integration for Image to Image #42zai-org/GLM-5.2discussion Tracking integration for Table to Text generation #47Request
Could you please add
IntelligenceLab/Long-Horizon-Terminal-Benchto the benchmark allow-list so that the Benchmark badge and leaderboard can render on the dataset page?Happy to provide any additional information if helpful. Thanks very much!
cc @julien-c @NathanHB