If It Ain't Broke, Don't Fix It: Failures of Epistemic Control in Language-Model Agents
Abstract
Language-model agents increasingly change real systems in response to test reports, execution traces, and messages about what changed. Before acting, such an agent must decide which explanation of the evidence to believe, how much of the system to change, and whether the evidence can be trusted. Existing evaluations score the final outcome or one of these decisions in isolation, so they cannot show where the path from evidence to action breaks. We introduce an executable testbed, covering discrete controllers in operational settings, feedback control of a continuous dynamical system, and Python programs, in which the correct answer is known exactly and every revision is executed. In the controller tasks, the environments the evidence allows are enumerated, every revision runs in all of them, and a model's stated belief can be replaced while everything else stays fixed. With two hidden changes, each model keeps the true environment among its plausible hypotheses in 80.0–97.8% of tasks, yet the leading hypothesis is correct in only 5.8%, below the 27.8% expected by chance. Replacing a wrong leading hypothesis with the truth raises selective repair (fixing what changed and nothing else) from 3.3% to 35.4%, so the stated belief guides action; yet even when told exactly what changed, eight of ten models repair selectively on fewer than half of their tasks. Models also show weak epistemic control, acting on reports without checking them: relative to a truthful passing report, a fabricated failure report raises edits to correct programs by 84.5 percentage points, and six of eight models never request an available test when a fabricated report says faulty code passes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.