Restoring Inputs Is Not Restoring Learning: Sparse Restoration in Multimodal Training
Abstract
Training on reduced visual views saves computation, yet restoring full images at evaluation need not recover the capability lost during training. Controlled continuations locate a recoverable contribution in the full-minus-reduced gradient at a shared learner state, beyond extra-compute and extra-update controls. Equal-norm remaining errors nevertheless produce opposite short-horizon loss changes. A local finite-horizon expansion explains this distinction through error transport across parameters and optimizer memory, followed by a signed terminal-task projection. GRAFT-V therefore restores the current raw reference-gradient mean with sparse same-state pairs and inverse-probability weighting before one optimizer update, rather than predicting a future task direction. Recovery repeats across six learner configurations; on the primary five-task experiment, fixed sampling closes 92.6% of the 2.44-point gap, while domain-stratified sampling closes 95.9% using 30.8% fewer final-training GPU-hours than reference training. The stronger-view and short-budget comparisons locate the operating boundary: preserving useful evidence directly can cost less than restoring its learning signal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.