acceptodds
Under review as a conference paper at ICLR 2027

Inference for Transformer Representation Shifts Through Metric Fields

Abstract

How can we determine whether a transformer has undergone a systematic shift in its internal representations? Existing methods typically summarize activations over a fixed collection of inputs, but do not directly quantify whether an observed difference persists across a population of input sequences. We formulate representation-shift detection as a population-level inference problem. Each randomly sampled sequence induces a hidden-state surface over layer depth and token position, which we treat as a random functional object and characterize through its induced metric tensor field. This representation is invariant to fixed global orthogonal transformations while preserving the spatial organization of model representations. Using paired inputs, we define the expected difference between metric fields over a prespecified task distribution as the inferential target and develop a global spectral test for detecting systematic shifts. In a controlled study of continued training, replacing only 1% of the training corpus with data from a different source produces a detection signal more than two orders of magnitude larger than that under the batch-order control. We further introduce a tolerance-based screening rule to distinguish small but detectable departures, such as those arising from uncontrollable noise, from substantively larger representation shifts. Together, these results provide a statistical framework for detecting, quantifying, and spatially characterizing transformer representation shifts at the population level.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.