acceptodds
Under review as a conference paper at ICLR 2027

Can Probes Reliably Guide Agent Decisions? Evidence from Controlled File Workflows

Abstract

Can a probe trained on a language model's internal state reliably guide an agent's next action? Not necessarily, even when the task state remains readable. We study a file task: test the current contents, or publish them if a successful test already applies. This setting lets us change the visible history without changing the correct decision. In a confirmation on 32 new file-roster families, testing another file reduces a fixed, feedback-trained Qwen probe's pairwise accuracy from 0.984 to 0.428, while the model's output scores retain 0.988. Both prespecified tests pass Holm correction. An exploratory centroid trained only on initial histories achieves 0.996 at the same layer and token position: the distinction remains linearly accessible there. Ranking is not enough, however: with the added test, the output score classifies only 53.9 percent of inputs correctly at its old threshold. Workflow experiments in Qwen and Gemma show why the reuse question matters: feedback training repairs repeated decisions, but the Qwen confirmation exposes a limit of that repair. Readable task state need not yield a reusable decision rule. The controlled task separates failures of ranking, thresholds, and feedback handling; it does not establish their frequency in real software projects.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.