One Bit in a Hundred: The Sparse Signed Substructure of Behavioral Weight Edits
Abstract
Behavioral weight edits, such as task-vector negation, debiasing deltas, and persona or refusal edits, are usually treated as dense, full-precision objects that are trained and then merged. We ask how much of such an edit is actually doing the work. Using social-bias removal as a testbed, we train low-rank contrast vectors from a model's own likelihood-based self-diagnosis and then strip them to one bit per weight and 1% of coordinates. Under a deployment-honest collateral budget in which every metric is measured with the intervention active, binarization retains 87% to 105% of full-precision removal across models from 2.6B to 32B, and a seven-condition factorial shows where the information lives. The true sign field at magnitude-selected coordinates recovers +0.353 of bias while the same field at random coordinates recovers +0.070; sign permutation, layer shuffling, and bottom-magnitude support collapse to ≈ 0. On a fixed grid of edit scales, the selection effect is +0.284 [+0.169, +0.407] on the ten pre-registered cells, unanimous, and +0.201 [+0.125, +0.284] on a twenty-cell scale-up spanning four benchmark families, positive in 19 of 20 cells. The concentration is a property of the edit scale. An adaptive, bias-blind rule that sets each edit's scale from a held-out collateral budget finds that random-coordinate signs tolerate about four times the scale of magnitude-selected ones, and at those scales the gap between them shrinks to intervals covering zero at both tiers tested; magnitude selection therefore buys the effect at a fraction of the scale and collateral. The resulting edit merges into the checkpoint as a persistent intervention, with no serving-time hook and no additional inference-time compute. At matched collateral and selection space its removal is statistically tied with activation steering and SentenceDebias driven by the same signal, DPO trained on that signal removes no detectable bias, and prompting fails the budget in 16 of 18 conditions, so what distinguishes the edit is persistence rather than raw efficacy. Alongside the method we contribute an evaluation protocol built on integer-item budgets, always-on collateral measurement, equal selection spaces, and per-method nulls. We also characterize a regime of low pre-existing bias in which the edit backfires and show that an abstention gate built on it fails held-out validation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.