Fragile Follow-Through: Uncovering How User Revisions Disrupt Long-Horizon Task Execution in LLM Agents
Abstract
LLM agents are increasingly entrusted with complex, long-horizon tasks, where fulfilling a user's request requires coordinating decisions and actions across many interdependent steps. Existing work typically evaluates agents against requirements specified at the outset, although recent studies have begun to consider user interruptions and revisions during execution. However, current evaluations do not adequately capture revisions prompted by task progress or explain how they disrupt long-horizon task completion. To bridge this gap, we adapt long-horizon tasks by altering their initial instructions and introducing mid-execution revisions that restore the original requirements. This yields ReQuest, an application-grounded benchmark designed to evaluate task completion and analyze agent behavior under dynamic user revisions. Revisions are triggered following observable progress under the altered requirements, with the agent's history and environment state preserved. Extensive experiments show that user revisions reduce average task success rates by approximately 17–63% relative to the original tasks across models. Process analyses help explain these losses: agents retain incompatible prior decisions, lose valid progress, or leave unchanged requirements unmet. Answers and guidance targeting these problems yield uneven benefits, pointing to mixed difficulties behind fragile follow-through in long-horizon tasks. Our benchmark, code, and evaluation protocols will be publicly available to support reproducible evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.