acceptodds
Under review as a conference paper at ICLR 2027

Representations in Motion: Layer-wise Rotation and Trajectory Divergence in LLMs

Abstract

Recent work suggests that activation dynamics–how representations change during inference–might be more revealing about language model behavior than static readouts from individual layers. In this work, we investigate the layer-wise evolution of representations of factuality using true and false statements when the model is prompted to either lie or to be honest. We report a geometric phenomenon in which the factuality directions extracted from the lying and honest conditions appear to rotate away from each other across layers. We find that such layer-wise rotation also occurs in a task that requires the model to take different visual perspectives in the 3D world, suggesting that the rotation phenomenon might be indicative of manipulating complex conceptual spaces in general, rather than unique to factuality. We then investigate the "real time" transformation of conceptual representations on a trial-to-trial basis. We show that it is possible to make strong predictions about the model's target task and behavior from trajectories alone. In addition, diverging points in single-trial trajectories correspond to the onset of rotation in aggregated-trial analysis. Finally, we show that both rotation and divergent trajectories occur in more naturalistic settings in which models are not overtly prompted to "lie" but rather report false statements in the course of completing a different task.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.