acceptodds
Under review as a conference paper at ICLR 2027

WorldModelLens: From Attention to Rollouts - A framework for World Model Interpretability

Abstract

Interpreting world models requires action-, step-, and rollout-level abstractions that existing language-model hook-and-cache tools don't provide, forcing researchers to rebuild bespoke analysis code for every new architecture. We present WorldModelLens, an open-source toolkit that unifies recurrent state-space models (e.g., Dreamer), token-based models (e.g., IRIS), and joint-embedding architectures (e.g., I-JEPA) behind a single typed adapter API and a shared hook-and-cache layer, so researchers can write one analysis and run it across model families. Applying it to the I-JEPA ViT-H/14 checkpoint, we find that Integrated Gradients achieves a small but statistically robust localization signal in insertion-based patch restoration (-0.19 standardized effect, p < 10⁻⁴), a much weaker version of the effect found in pixel-reconstruction masked autoencoders (-1.15), suggesting I-JEPA's predictor relies more heavily on distributed context than a pixel-reconstruction objective does, though not to the point of being fully non-localizing. Our latent-attribution method recovers this same ranking while requiring  180× fewer model evaluations than exhaustive patching, and correlates with costlier head and block-level interventions (r=0.70-0.93). By collapsing architecture-specific plumbing into a common adapter surface, WorldModelLens lets the same hooks trace a prediction from an individual attention head up through block-level dynamics to full rollout behavior, turning interpretability of world models from a per-architecture engineering problem into a matter of choosing where, on that continuum, to look.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.