Local Substitutions Do Not Compose
Abstract
A model can be nearly insensitive to individual identity substitutions while still depending on the correspondence among many inputs. When those identities are forgotten jointly, the resulting behavioral loss need not remain small. We formalize this distinction through behavioral quotients, which identify concrete states under declared transformations and measure the irreducible prediction error induced by that abstraction. Under uniform coalition queries, we derive sharp bounds on block quotient error from singleton contextual substitutions and finite query certificates that remain valid after the quotient is selected from the same measurements. We then show a fundamental failure of composing individually safe substitutions: every singleton discrepancy can vanish as the number of bound objects grows, while any accurate block quotient still requires states at a fixed positive error tolerance. Coordinated renaming that preserves the bindings instead gives an exact quotient with states. Controlled Transformer experiments reproduce the corresponding finite-size separation in learned responses. Finally, we decompose shared prediction error across aligned instances into residual identity dependence, cross-instance disagreement, and fitting error, with direct diagnostics for the first two terms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.