acceptodds
Under review as a conference paper at ICLR 2027

Agents Represent Action Reversibility, But Rarely Say So

Abstract

Models can tell an irreversible tool call from a recoverable one. Under a neutral system prompt, a mean of 91.9% of the irreversible ones draw no caution at all on the external set. Reversibility—whether an action is idempotent, reversible, compensable, or irreversible—is linearly decodable from a model's internal activations in all 19 open-weight models we test, above an action-name baseline on our internal set and in 13 of 18 models on externally authored schemas. The silence is elicitation rather than ignorance: telling the model to flag risky steps cuts it, on our internal set, from 84.9% to 27.9% without closing it, and a generic instruction to take care closes almost none of it. Asking outright does not fix it either. Models assert that an irreversible action can be undone in 26% of the judgments they make. On our internal set a linear probe's scores rank an irreversible action above a recoverable one 86.9% of the time, and tell a right spoken judgment from a wrong one 53.9% of the time, where 50% is chance. The same dissociation appears in causal control. Steering works in seven of ten models when the direction is the gap between the two classes' average activations, and in one when it is our probe's own. What decides this is how the direction is fitted, not which model it is—our probe ranks actions well, but two fits on different halves point different ways. The direction that decodes best is close to the worst one to steer with. The readout is useful in its own right. The probe raises detection F1 over the model's own caution language by 1.70× across eleven instruct models, and beats a prompted classifier. It transfers across models and to schemas written outside this work, though not to pooled live decoding, and under a threshold fixed in advance it clears about 30% of our internal action set (23% on a current scikit-learn release) while letting through only a few percent of the irreversible ones, at negligible latency. We release the benchmark, the probes, and our evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.