The Reference Matters: Displacement Probing for Weight-Space Learning
Abstract
Weight-space learning predicts a model's attributes directly from its parameters, and probing-based methods do so efficiently by reading a model's responses to learned probes. Prior work shows that probing succeeds within a Model Tree—models fine-tuned from a shared pretrained origin—but not why. We trace this to the shared origin itself: since a tree's leaves start from the same weights, what distinguishes them is their displacement from that origin. This motivates probing the displacement rather than the raw weight, an intervention we call the DiSplacement Probe (DSProbe). We develop this idea along two axes: extending it to second-order Gram responses, whose displacement we show encodes the learned features induced by fine-tuning; and treating the reference as a design choice, learning a shared component to subtract beyond the pretrained origin. Across four architectures on the Model Jungle benchmark, the resulting Refined-DSProbe (R-DSProbe) sets a new state of the art, matches prior accuracy with substantially fewer probes, and supports a single encoder shared across Model Trees, reducing metanetwork parameters by 3.61 times.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.