LongJudge-Bench: A Long-Horizon Agent Judge Benchmark for Intervention Decisions
Abstract
Long-horizon LLM agents execute stateful workflows in which tool actions alter persistent environments and determine delayed outcomes. Supervisors must therefore choose what to do before completion, yet existing judge benchmarks mainly score completed responses or trajectories and do not directly evaluate which executable intervention best preserves future success. We introduce LongJudge-Bench, a consequence-grounded benchmark with 70 unfinished checkpoints from 35 AppWorld tasks, balanced across five domains and seven intervention families. For each case, we execute all seven actions on five matched frozen scenario–seed realizations under a shared executor, checker, budget, and continuation contract, producing 490 empirical action values from 2,450 outcomes. Judges select from pre-outcome evidence, and we separately measure response validity, retained task success (TS), exact reference-action accuracy (AA), and regret. Across 13 models under the Original interface, 90.22% of responses are valid, but mean TS is 53.47% and mean AA is 36.37%, exposing distinct gaps in consequence preservation and reference recovery. Under Declared semantics, observed TS and AA rise for all six tested models, while history and position controls show that AA is consistently more sensitive than TS to evidence representation. These results distinguish valid outputs, consequence-preserving choices, and exact reference recovery as separate capabilities. provides a controlled basis for evaluating actionable supervision in long-horizon agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.