acceptodds
Under review as a conference paper at ICLR 2027

When Historical Energy Is Insufficient for Local Training Response

Abstract

A training trajectory can be nearly one-dimensional in historical energy, yet discarding its weak second direction can lose an order of local prediction accuracy. For smooth fixed-loss gradient descent over fixed finite windows, we characterize the necessary and sufficient gradient and curvature capture rates for preserving the unrestricted linearized model's normalized parameter-error guarantee. With nonzero gradient and transverse curvature, even the best line through the anchor, fixed across the window and chosen using the complete future, has sharp error. Both the gradient-curvature plane and the exact noiseless historical top-two space attain error. History therefore contains a sufficient response space, but every fixed explained-energy threshold below 100% eventually discards its necessary second direction, whose energy fraction vanishes quadratically. Multi-scale tests in selected LoRA coordinates of eight LLM model-data blocks support the predicted orders and assess the sharp best-line coefficient, with model-dependent finite-scale deviations. Across four families, matched experiments locate over 96% of the one-direction predictor's squared residual in a direction carrying under 2% of future update energy (family medians). A varying-driver motion law separates driver and curvature contributions to historical compression; shuffled-AdamW diagnostics examine direction selection beyond fixed-loss GD.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.