acceptodds
Under review as a conference paper at ICLR 2027

What Stochastic Information Matters for Learning? Observable-Relative Equivalence and Control

Abstract

Two stochastic learning rules can have the same conditional update mean and full covariance at every common learner state yet lead to different finite-horizon outcomes. We establish that the stochastic information needed to describe learning depends on the future quantity being asked about. We characterize when a summary of today’s update is sufficient to predict a future learning outcome. For every finite order , we construct SGD systems on the same smooth, strongly convex objective whose temporal polynomial statistics agree through order , while their probabilities of the same finite-horizon event differ by arbitrarily close to one. No fixed finite moment hierarchy is universally sufficient. When summary-equivalent laws are realizable choices, unresolved stochastic structure becomes a control variable that can improve learning. At the same learner state, the law minimizing immediate error differs from the law minimizing error after eight further SGD steps; choosing for future loss reduces terminal risk by 8.35% relative to state-aware immediate-loss control in our exact system. In three neural settings, Gaussianizing updates while preserving their conditional mean and covariance moves boundary-crossing risk toward the Gaussian prediction in 59 of 60 systems, closing 59% to 75% of the discrepancy. In closed-loop on-policy learning, matched local mean and covariance lead to different future evidence distributions. In same-prompt Qwen RLVR, moment-matched gradient laws produce substantially distinguishable applied updates after clipping and AdamW. The future learning question determines both the stochastic resolution needed for prediction and the realizable law worth choosing for control.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.