Loss-Invariant Projections as Passive Probes of Learned Representations
Abstract
Learned feature representations in neural networks often contain structure beyond that directly used by the final task output. We study such structure using passive probes: fixed, untrained, property-independent projections that remain unchanged while the representation evolves during training. We motivate this approach through prediction on , where equivalent vector and Hermitian parameterizations reveal an additional loss-invariant trace coordinate. This motivates a general construction in which fixed random projections serve as observers of learned features. Because the observer is loss-invariant and independent of the property being studied, changes in accessibility reflect changes in the representation relative to the fixed observer rather than adaptation of the observer. We show that ensembles of passive probes can have direct relationships to task-relevant information such as target alignment. Across surface-normal estimation, image classification, and image inpainting, we observe different changes in accessibility under our constructions: eventual difficulty becomes increasingly accessible in surface-normal estimation and image inpainting which are regression tasks, whereas its accessibility remains near its initial level in image classification. Comparisons with learned linear probes further show that recoverability and passive accessibility can evolve differently during training. These results show how passive probes can separately characterize changes in representation geometry and the accessibility of eventual task difficulty.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.