acceptodds
Under review as a conference paper at ICLR 2027

Fragile Follow-Through: Uncovering How User Revisions Disrupt Long-Horizon Task Execution in LLM Agents

Abstract

LLM agents are increasingly entrusted with complex, long-horizon tasks, where fulfilling a user's request requires coordinating decisions and actions across many interdependent steps. Existing work typically evaluates agents against requirements specified at the outset, although recent studies have begun to consider user interruptions and revisions during execution. However, current evaluations do not adequately capture revisions prompted by task progress or explain how they disrupt long-horizon task completion. To bridge this gap, we adapt long-horizon tasks by altering their initial instructions and introducing mid-execution revisions that restore the original requirements. This yields ReQuest, an application-grounded benchmark designed to evaluate task completion and analyze agent behavior under dynamic user revisions. Revisions are triggered following observable progress under the altered requirements, with the agent's history and environment state preserved. Extensive experiments show that user revisions reduce average task success rates by approximately 17–63% relative to the original tasks across models. Process analyses help explain these losses: agents retain incompatible prior decisions, lose valid progress, or leave unchanged requirements unmet. Answers and guidance targeting these problems yield uneven benefits, pointing to mixed difficulties behind fragile follow-through in long-horizon tasks. Our benchmark, code, and evaluation protocols will be publicly available to support reproducible evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.