acceptodds
Under review as a conference paper at ICLR 2027

LongTool-Bench: Benchmarking Lifecycle-Aware Adaptation in Long-Horizon Tool Agents

Abstract

Long-horizon tool agents execute interdependent actions that modify persistent application state, often across applications. Yet tool capabilities, interfaces, permission boundaries, execution conditions, or identities may change after consequential effects have been committed. Successful adaptation must revise invalid assumptions while preserving valid progress, avoiding duplicate effects, and satisfying revised constraints. Existing benchmarks study long-horizon execution, stateful interaction, or unreliable tools, but rarely combine controlled in-episode evolution, retained side effects, and executable recovery obligations within an unfinished task. We introduce LongTool-Bench, comprising 200 AppWorld tasks, 400 evolution contracts, and 38,400 episodes across sixteen models. It injects changes at semantically eligible boundaries, preserves committed state, and varies access to change information while separately measuring task correctness, event exposure, contract recovery, prohibited behavior, recovery effort, and declaration reliability. Model-independent executable witnesses establish chain feasibility from retained boundary states. Relative to Hidden, Oracle-Docs improves event-chain resolution by 23.9 pp but Official Success by only 5.0 pp; moreover, 86.8% of chain-resolved episodes still fail the original task. Category analyses and worker ablations further separate failures in exposure, recovery, state protection, and terminal completion. These results show that contract recovery alone does not ensure end-to-end adaptation: dependable agents must jointly track evolving contracts, committed state, unfinished dependencies, and terminal task evidence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.