Probing Under Distribution Shift: Learning Generalizable Distributed Representations of LLM Activations
Abstract
Probing is one of the most promising approaches for monitoring the internal states of large language models: probes read a subset of model activations and predict an external variable of interest. Yet recent work has documented poor out-of-distribution generalization, overfitting. These are standard supervised-learning generalization problems, of which probing is a special case, with the additional challenge that its inputs come from the high-dimensional and highly structured activation space of a pretrained model. Therefore, we propose to focus on generalization and evaluate probes under distribution shift using **leave-one-environment-out evaluation**. This perspective suggests two natural independent methodological directions: distributionally robust learning that explicitly targets generalization across environments, and learning distributed representations of the activation space that best support generalizable predictions. We systematically compare baseline probes with distributed representations trained with and without distributionally robust objectives. Across three tasks and 15 environments, our proposed distributed representations consistently improve OOD generalization despite greater complexity, while none of the existing distributionally robust methods we evaluate provides additional gains. Our proposed distributed representation compresses the full activation tensor via two-stage inducing-point attention with a sparse layer selector, meaning that once it is learned the probe remain resource-efficient. Our results suggest that regularization is better achieved by training across diverse environments than by simply constraining probe complexity. Finally, we discuss implication for future research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.