acceptodds
Under review as a conference paper at ICLR 2027

Mitigating Agentic Misalignment via Jacobian Workspace Self-Supervision

Abstract

We show that using an AI agent's Jacobian workspace can mitigate agentic misalignment even after the agent begins rationalizing harmful behavior, substantially reducing harmful actions without retraining the underlying model. The Jacobian lens measures how changes in internal states propagate toward output representations, providing a basis for detecting when reasoning is heading toward a harmful action. We combine measurements derived from this lens with ordinary internal states in a small trained detector that monitors the agent as it reasons. When the detector's score crosses a calibrated threshold, we interrupt generation, discard the current reasoning, and ask the same model to write a brief corrective reminder. The agent then restarts from the task and reminder. An optional intervention steers the advisor's internal states while it writes the reminder. On a held-out i.i.d. evaluation set, our best interventions reduce agentic misalignment from to on Qwen3-32B and from to on Qwen3-14B. These results establish internal monitoring and guided self-correction as a promising complement to training-time safety interventions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.