acceptodds
Under review as a conference paper at ICLR 2027

The Exact Exchange Rate of Gradient-Free Credit

Abstract

Gradient-free training is back at LLM scale, and its cost is now the central question. Folklore holds that on a layer of parameters, scalar-broadcast evolution strategies need times more samples than backpropagation — directionally right, but the constant has never been measured. We derive it exactly: for antithetic rank-1 perturbations of an layer, the estimate from forward evaluations against the true gradient obeys with , parameter-free, exact in the linear regime at frozen weights; dynamic training adds landscape-dependent excess variance atop this floor. We verify it on every layer of the reference MLP (ratios 0.93–1.12) and of a frozen extreme-aspect-ratio chain (0.976–1.006), killing all four alternatives in that pre-registered experiment (maximum deviation 0.006 across a span of ). The law then prices, by extrapolation, the largest gradient-free training system to date: it operates at per-step and finances that signal with coherently integrated steps at the GPU-hours of backpropagation. Channel by channel, our measurements yield a design-space map (two primitives, four laws, six measured rows) and three impossibility statements, among them: no scalar-broadcast channel escapes the exchange rate, and approximate backpropagation is the wrong objective. This is a measurement study, not a new method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.