acceptodds
Under review as a conference paper at ICLR 2027

High-probability guarantees for linear accessibility in feature superposition

Abstract

Neural networks can leverage feature superposition to encode more concepts than dimensions, but cross-feature interference constrains the linear accessibility of simultaneously active features. By framing linear accessibility as a compressed sensing problem, we derive high-probability recovery bounds for fixed supports under subgaussian noise, proving that dimensions suffice for the per-input regime in which networks typically operate in, in contrast to the quadratic scaling required by uniform worst-case guarantees. We characterize the asymmetry between active and inactive interference and the trade-off between interference and observation-noise budgets. We then validate these bounds across system parameters through numerical experiments based on Gaussian-tail approximations. We also introduce IHT-SAE, which uses learned iterative refinement to improve feature recovery beyond the limits of linear accessibility. These results quantify the geometric constraints of the linear representation hypothesis, providing a framework for evaluating sparse autoencoders, compositional generalization, and neural interpretability. Code to reproduce this work is available at: https://anonymous.4open.science/r/Linear-Availability-07CD

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.