When Does Covariance Minimize the Expected Maximum Edit Response?
Abstract
Covariance-based knowledge editors choose an update direction by minimizing its mean squared response on unrelated inputs. A small mean square can hide a large response among the probes that evaluate the edit. We ask when the same direction also minimizes the expected largest absolute response over a finite probe budget. In a fixed feasible subspace, covariance placement is optimal for every nonzero target and every number of iid probes exactly when the symmetrized whitened activation law is spherical. Centered elliptical laws with finite covariance meet this condition, including laws with polynomial tails. Bounded laws with the same covariance can need different directions at different budgets. A dimension-free converse turns measured directional anisotropy into a lower bound on the worst-target risk ratio. The remaining gain depends on the target. In 32-dimensional target-conditioned slices from five architectures, the median held-out risk ratio of covariance placement to a fitted alternative at 128 probes is 2.20 at searched worst targets and 1.048 at real edit targets. On ROME's edits in two architectures, budget-fitted placement replaces the quadratic objective. It has lower geometric-mean held-out risk than ROME and a matched quadratic refit at every evaluated budget at or above its fitting budget, with rewrite accuracy unchanged. For response magnitude, fitting an empirical Weibull shape parameter raises the held-out coverage of the fitted sub-Weibull predictor from 90.1% to 97.9% at 5,000 probes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.