acceptodds
Under review as a conference paper at ICLR 2027

Dynamic Mechanistic Interpretability for Real-Time State Transitions in World Models: Tracking Thoughts in Motions

Abstract

Mechanistic interpretability has made impressive strides in opening up the "black box" of neural networks, but existing methods like sparse autoencoders and activation patching share a fundamental limitation: they treat models as static snapshots. This creates a major blind spot when analyzing world models, such as JEPAs or model-based RL agents—whose internal representations must dynamically evolve as environments shift. In this paper, we introduce dynamic mechanistic interpretability (DMI), a novel framework built to track, isolate, and causally verify time-varying computational circuits in real time. At the core of our approach are ‘temporal sparse autoencoders’ (t-SAEs), which enforce smoothness constraints across consecutive time steps to extract persistent, interpretable primitive features. Building on this, our dynamic circuit discovery (DCD) algorithm applies sliding-window causal patching to continuously map how computational subgraphs reconfigure during online adaptation. We evaluate DMI on world models operating across non-stationary continuous control environments. Our findings demonstrate that DMI isolates state-dependent sub-networks using under 15% of total edge weights while outperforming static baselines in reconstruction faithfulness. Finally, we show that targeted real-time interventions on these dynamic circuits allow us to directly manipulate an agent’s internal belief state. By moving interpretability from static snapshots to dynamic systems, DMI provides a missing blueprint for understanding and controlling adaptive AI.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.