Agents Verify Their Assumptions, Not the Specification: Interpretation Lock-in in Autonomous Coding Agents
Abstract
Large language models (LLMs) have become capable coding agents that complete long software tasks autonomously and test their own work. While LLMs excel at well-defined programming tasks such as code generation and bug fixing, their self-verification breaks down when a requirement is ambiguous, because the tests they write share the interpretation of the code they test. In this paper, we study this failure, which we call interpretation lock-in: agents commit to one reading of an ambiguous requirement, implement it, and validate their work against that same reading. In a controlled suite where every valid reading is known, and in pipelines rebuilt from long-horizon benchmark runs, the resulting programs crash or silently produce wrong output under other valid readings. The failure is largely one of commitment rather than knowledge: asked only to list plausible readings, the same models often name the one they missed. To remedy this failure, we propose reading-diverse self-training (RDST), which rewards an agent’s own so- lutions for the number of valid readings they handle and fine-tunes the agent on its best episodes. Experimental results show that, relative to rewarding success on a single reading, RDST substantially reduces lock-in, including on task families held out from training, at little cost on well-specified tasks, and that the behaviour carries over to the interaction protocol of another agent harness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.