acceptodds
Under review as a conference paper at ICLR 2027

Second Order Delta Space Search (SODSS),

Abstract

Which weights of a trained network can be pruned or quantised is usu- ally decided from the loss landscape at convergence, using gradients on calibration data or extra Hessian-vector products. We show that the train- ing trajectory already contains this information and that it can be read from checkpoints alone. Along a gradient-descent path the second differ- ence of the parameters satisfies an exact identity, at =−η Htvt−1, with a path-averaged Hessian. We characterise what this gives: the curvature the optimiser experiences along its own path (a Rayleigh quotient), glob- ally and per block, with a bias under minibatch noise bounded by 1/(η∆) that vanishes with the checkpoint stride ∆; and what it does not: the Hessian diagonal per parameter, which we show is not identifiable. Second- Order Delta-Space Saliency (SODSS) builds on this. It detects second- order events in parameter, representation and token-embedding space with a scale-free statistic (the second difference of log-speed), segments training into temporal regimes, and scores each weight by a regime-weighted path- Fisher saliency θ2 N,j t wtv2 t,j , which we prove is an Optimal-Brain-Damage saliency with the empirical Fisher averaged over the path and which reduces to Synaptic Intelligence without the θ2 factor. The score needs no gradients, data or Hessian products and is computed in one streaming pass with O(d) memory. On SGD-trained networks (an MLP and a CNN) one-shot global pruning with SODSS is within a point of magnitude pruning up to 95% sparsity, beats gradient-based Fisher and SNIP criteria throughout, and is clearly best at 98%; the event detector recovers learning-rate schedule changes and representational phase changes without being told the sched- ule. Under AdamW the raw score fails, because the premise vt ∝gt does not hold; undoing the preconditioner with the second-moment state that checkpoints already contain restores it, and the corrected score then beats magnitude pruning on an AdamW-trained Transformer at every sparsity. We also report what does not work: per-parameter second-order statistics do not help, and regime weighting has no measurable effect.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.