Gradient Descent on Tokens Is Sparse: Understanding the Computational Structure of Differentiable Token Optimization
Abstract
Inference-time computation has become an important factor in language model quality, and differentiable token optimization (DTO) is one way to allocate this computation. DTO treats a model's discrete output tokens as continuous parameters and refines them using gradient descent. Test-time reasoning, adversarial attacks and controllable generation all share this template, and every member's default gives each position the same fixed budget at every step, presuming that the computation is homogeneous. This assumption does not hold, because DTO is a discrete process evaluated through a continuous surrogate, and the two decouple within the first few dozen updates. Once a position's margin exceeds what a single step can overturn, the surrogate's gradients can no longer alter the discrete output. The resulting computation is not merely redundant, since updates to uncommitted positions can continue to alter the output without improving quality. Accuracy falls monotonically from 47.1% at greedy decoding to 34.2% at T=60, with losses outnumbering gains 32 to 1. Because the decoupling arises from the interface between continuous optimization and discrete token assignment rather than from any particular objective, it holds across the family and is measurable. Tracking 340 discrete trajectories over two model families on MATH-500, we find at least 89% of token flips within the first 50 of 59 measurements, the most active fifth of positions absorbing a third of the churn, and a churn plateau that never converts into quality. Three propositions characterize what can be recovered from the decoupling and yield three mechanisms, gradient reuse across stable updates, position-level activity prediction, and objective-driven stopping with a loss bound. Acting on the structure jointly removes 15.8% of model calls at T=60 without measurable loss of accuracy. The budget of DTO is therefore a measurable property of the decoupling rather than a constant to be tuned.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.