acceptodds
Under review as a conference paper at ICLR 2027

Where Models Diverge: Linear Predictability Trace the Direction of Representational Change

Abstract

We study what directional differences in cross-model predictability can reveal about how two models organize their representations on the same data. To do this, we fit linear maps between a pair of models' representations in both directions and use their predictive asymmetry – the difference in held-out linear predictability between models – as a diagnostic tool. This signal reveals which distinctions one model preserves relative to another, localizes which examples drive these differences, and can identify where relative performance gains between two models lie, without requiring any labels, concept/dictionary extraction, or clustering. Validating this diagnostic through controlled experiments, we show that it can recover differences in supervision granularity, training data coverage, and training progression between model pairs. Changing only the supervision objective can even reverse the direction of linear recoverability between a fine-tuned model and its parent. We apply predictive asymmetry to twelve domain-specialized foundation models and trace how their representations differ from their parent models after fine-tuning. Our experiments demonstrate that predictive asymmetry gives a label-free account of specialization: it tracks relative training exposure between specialist and parent and localizes where specialist–parent probe advantages hold. Predictive asymmetry therefore provides a simple, label-free view of where and in which direction model representations diverge.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.