Competitive Repair Capacity: When Valid Neural Edits Stop Competing
Abstract
Model editing certifies that repair is possible with geometric quantities: editable rank, null- space dimension, codebook size. It leaves a second question open: an acceptable, successful edit can exist and still lose to an unacceptable alternative under the objective that selects edits. We define competitive repair capacity, the number qof valid successful actions beating every unacceptable one, and characterise it exactly: in finite families qis the deletion distance to optimizer-facing failure, its margin-refined form sharply characterises robustness to simultaneous action loss and bounded score error, and continuous families obey a volume law in which admissible dimension alone cannot fix capacity. Two real systems show these are not redundant. On a published GPT- 2 XL AlphaEdit locus competitive capacity collapses to zero across a one-dimensional increase in null-space capacity; one GRACE reproduction is aligned at its released radius, fragile under an off-default geometry stress. An equation- and code-level audit of 41 protocols finds none exposing competitive capacity. A theorem-derived certificate then isolates aligned decisions prospectively — zero observed competitive failures among 141 certified states — yet abstains too often to deploy. Together these results separate the existence of repair capacity from its competitive, certifiable and deployable forms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.