acceptodds
Under review as a conference paper at ICLR 2027

BEYOND VARIANCE MATCHING: DATA-DEPENDENT LAYERWISE STARTS FOR DIFFUSION NETWORKS

Abstract

Diffusion networks are almost always started from variance-preserving random draws such as Xavier or He initialization and then refined by an optimizer. These initializers stabilize activation and gradient scales, yet they ignore the training dis- tribution: the first gradient steps must discover task structure that a cheap linear algebra pass could already suggest. We ask whether a data-dependent start can place ε-prediction diffusion denoisers at better weights that reduce subsequent training cost. We introduce Layerwise Closed-Form initialization (LCF), which visits each affine map in depth order, fits a throwaway linear probe to form a residual-augmented target, takes a single exact line-search gradient step from the layer’s default random draw toward that target, restores the true nonlinear opera- tor, and proceeds to the next layer. Each hidden weight is thus fit by least squares on the post-activation features of the layers below it, rather than obtained by fac- torizing one linear map across depth. On 12 diffusion datasets × 5 denoisers, LCF improves step-0 validation error on 60/60 settings; on convolutional denoisers it also yields positive net optimizer savings after subtracting its own cost, on 12/12 large diffusion CNNs (mean +19.7 net steps) and 6/6 base CNNs on grayscale sets (mean +33.5). At this loose target a zero-initialized output layer matches these savings at no cost; at targets within 1.1–1.5× of the converged loss, LCF saves more net steps than it on 9/12 datasets. These results support a training-cost savings claim for diffusion CNNs: a cheap closed-form start can replace optimizer steps.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.