acceptodds
Under review as a conference paper at ICLR 2027

Long-horizon-terminal-bench 2.0: Testing the limits of agents on long-horizon terminal tasks with dense reward-based grading

Abstract

AI agents have become increasingly capable of autonomously completing short and well-specified terminal tasks. However, existing terminal benchmarks either focus on short, simple problems that finish within a few minutes or on hard tasks evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, leading to sparse reward signals and an incomplete picture of agent capability. We curate Long-Horizon-Terminal-Bench 2.0 (LHTB 2.0), a diverse set of 77 long-horizon tasks grouped into 13 domains. Each task follows a Terminal-Bench-style Harbor setup with a reference solution or simulation engine and is further decomposed into fine-grained graded subtasks. We design dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on long-horizon workflows. We evaluate 15 frontier models on LHTB 2.0. The strongest current system, GPT-6-astra, reaches the soft-pass threshold (R ≥ 0.95) on 37/77 tasks (48.1%) and achieves a dense mean reward of 0.68, followed by GPT-5.6-sol at 32/77 tasks (41.6%) and a mean reward of 0.67. Across 1,145 scored task runs, only 16.8% pass the soft-pass threshold, while 15.4% make no meaningful progress (R < 0.05) and 67.9% achieve partial progress. These results reveal substantial headroom for improvement. We further analyze common failure modes and error patterns of models and agents on long-horizon terminal tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.