acceptodds
Under review as a conference paper at ICLR 2027

Representation Symmetries and Architectural Constraints on Probe Identifiability

Abstract

A central question in neural representation analysis is when observations of hidden representations uniquely determine an associated symbolic assignment. A symbolic label may be predicted accurately from a representation without being uniquely determined by the observed information. Architectural information can restrict hidden-coordinate changes that preserve a neural model and its observations, but invariance under known symmetries does not by itself imply identifiability. We distinguish observational identifiability, symmetry invariance, and agreement of fixed linear sign-probe predictions. A general observation-factorization criterion characterizes exact extractability, while full-space and boundary-local probe criteria identify the positive-ray condition for exact agreement. We then study finite representation calibration, where allowed orthogonal transformations fix specified hidden states pointwise. For an explicit norm-and-inner-product observation protocol on a known linear subspace, we classify every observational fibre and characterize both global and fibrewise probe identifiability. Under Gaussian evaluation, the largest disagreement over all calibration-preserving transformations has a closed form determined by the uncalibrated component of the effective probe and is attained by a Householder reflection. Independent Gaussian calibration gives an exact finite-sample law governed by the support dimension. Architectural stabilizers provide upper bounds, not automatic completeness certificates. Controlled FP64 experiments agree with the finite-calibration predictions. Calibration-preserving weight changes in Pythia-1.4B and Pythia-2.8B still change fixed-probe predictions, while complete-forward checks identify a reproducible numerical deviation under one joint intervention.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.