acceptodds
Under review as a conference paper at ICLR 2027

Loss-Invariant Projections as Passive Probes of Learned Representations

Abstract

Learned feature representations in neural networks often contain structure beyond that directly used by the final task output. We study such structure using passive probes: fixed, untrained, property-independent projections that remain unchanged while the representation evolves during training. We motivate this approach through prediction on , where equivalent vector and Hermitian parameterizations reveal an additional loss-invariant trace coordinate. This motivates a general construction in which fixed random projections serve as observers of learned features. Because the observer is loss-invariant and independent of the property being studied, changes in accessibility reflect changes in the representation relative to the fixed observer rather than adaptation of the observer. We show that ensembles of passive probes can have direct relationships to task-relevant information such as target alignment. Across surface-normal estimation, image classification, and image inpainting, we observe different changes in accessibility under our constructions: eventual difficulty becomes increasingly accessible in surface-normal estimation and image inpainting which are regression tasks, whereas its accessibility remains near its initial level in image classification. Comparisons with learned linear probes further show that recoverability and passive accessibility can evolve differently during training. These results show how passive probes can separately characterize changes in representation geometry and the accessibility of eventual task difficulty.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.