acceptodds
Under review as a conference paper at ICLR 2027

LLM Agents Mistake Success Elsewhere for Completion Here

Abstract

Testing operations in isolated environments before production execution is a standard safety practice. When an agent retains this test history, can it correctly determine what still needs execution in the target environment, even when the history explicitly identifies where each operation occurred? We evaluate 13 deployments on controlled transformations of 60 public tasks from the benchmark family and AppWorld, preserving native tools and authentic receipts. We assess actual effects with native state checks and completion claims with human-validated LLM annotation. With complete environment mappings recorded separately from receipts, successful rehearsal history raises false completion claim rates from 1.9% to 79.7% on and from 7.7% to 69.9% on AppWorld, compared with matched pre-execution controls. The increase occurs in all 13 deployments on both task sources. Copying the already available environment relation beside its receipt reduces these rates to 18.6% and 24.7%, respectively, and substantially improves completion of the required operation. Changing receipt formatting alone produces no comparable reduction. These results show that truthful history and complete execution-scope information do not ensure correct decisions about what still needs execution. Success is environment-specific: reliable continuation requires applying that scope to both actions and completion claims.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.