Shrinkage in Zeroth-Order Optimization: Estimation Risk and Expected Descent
Abstract
Reducing the estimation error of a zeroth-order (ZO) update need not improve optimization after its step size is chosen. We study this distinction for anchored ridge shrinkage under exact finite batches of random directions. Optimally rescaled Euclidean risk and optimized isotropic quadratic descent rank updates identically. For iid sphere updates, an explicit Stein kernel gives risk bounds at a specified nonzero shrinkage strength. With the direction count fixed at and dimension sufficiently large, the same zero-anchor rule has certified risk reduction and strictly worse optimized descent. We quantify this relative loss and give explicit finite-dimensional enclosures. The nonlinear coefficient remains random in the limit, making effective steps vary across updates. In a specified curvature family with bounded anchors, the rule can beat Raw while remaining asymptotically inferior to the fixed mixture . Independent simulations test these predictions. At frozen OPT-125M checkpoints, expanded step searches yield interior fitting-loss minima in all forty selections, yet the nonlinear rule has no consistent advantage. Independent profiling predicts some coefficient effects at small steps, with reduced predictive agreement at the full selected step. A local audit obtains informative risk and optimization bounds but no final endpoint confirmation in six budget–curvature comparisons. The results separate estimation risk, expected descent after choosing a step size, and statistical confirmation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.