Mask Faithfulness Does Not Guarantee Selective Control
Abstract
A parameter decomposition writes a weight matrix as a sum of rank-one components with learned gates. Its standard test is mask faithfulness. The network is expected to behave the same under every mask allowed by the gates. We explore a different question. Can a faithful decomposition erase one behavior while preserving others? We answer this exactly on a controlled task with independent, always-active behaviors. For every and any capacity , we show that every faithful decomposition with the fewest active gates pairs coordinates. This is because a paired read cancels on half the inputs, so it appears sparser than a single coordinate. Erasing one behavior then raises another behavior's error by at least (MSE), for any real coefficients, whereas the coordinate basis erases with no damage. The sparsity advantage that pairing creates depends on the input data. An input that activates only one behavior activates both the sum and difference reads in a pair, but only one coordinate read. If single-behavior inputs become frequent enough, the coordinate basis is the optimal choice. A second theorem extends this bound to smooth targets, small input noise, continuous gates, and a finite penalty. It holds for every -approximate optimum. In three small trained networks, the same change of training data cuts the provable cost of selective edits by more than . Therefore, selective control should be judged separately from masked reconstruction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.