Structurally Close, Security-Wise Far: Contradictory Policies in Neural Networks
Abstract
How close can two neural subnetworks be if they implement essentially the same benign functionality but contradictory security policies? We formalize this question through structural security distance (SSD), the minimum Hamming distance between masks of a common parent network that approximate a prescribed safe/unsafe policy pair. We also distinguish unrestricted SSD from distance under mask-size constraints and from the distance between a specified safe mask and the nearest unsafe realization, since these quantities need not agree. We establish four main results. First, structural proximity depends on which realizations are compared. We construct a perturbation-stable ReLU family in which the minimum-size safe and unsafe masks are \(2q\) edits apart, the minimum safe mask is \(q\) edits from the nearest unsafe realization, and the unrestricted SSD is a single mask-coordinate change. Thus, a globally close safe/unsafe pair does not imply that a particular deployed safe model is close to an unsafe realization. Second, low-dimensional safety localization does not imply a small number of structural edits: a one-dimensional safety direction can require \(\Theta(d)\) binary neuron-gate changes, and under additional assumptions on benign representations, every discrete one-gate-at-a-time transition can incur a constant benign-error barrier even though continuous directional ablation preserves benign behavior. Third, estimating normalized SSD remains NP-hard in the worst case even at fixed nonzero approximation tolerance, with guaranteed feasible safe and unsafe masks, known minimum sizes, bounded sparse inputs, and robustness to parameter perturbations. Finally, we show that any transition between opposite policies must accumulate sufficient functional change, yielding certificates for the safety of a deployed mask within a bounded number of edits.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.