acceptodds
Under review as a conference paper at ICLR 2027

How Reversible Normalization Changes Linear Forecasting

Abstract

We characterize how reversible normalization affects the function class, fitting weights, and inference-time behavior of a linear forecaster. Element-Wise Normalization (EWN) learns separate probability distributions over input lags for centering and scale estimation. These distributions are optimized by differentiating the forecasting loss through the ridge solver used to fit the final head. With a homogeneous head and shared input and output statistics, the learned scale cancels exactly at inference, even when it depends on the input. During fitting, however, squared residuals are weighted by the inverse squared scale. The learned center weights determine the affine set of attainable lag maps. Under our compressed representation, a center-weight vector outside the feature span adds one direction relative to instance normalization (IN). An intercept in normalized coordinates introduces a scale-dependent term that is generally nonlinear in the input. In a controlled heteroskedastic study, learning the scale improves estimation without changing the deployed function class. No clear advantage over IN is observed under the null condition with uniform lag weights, supporting an explanation based on residual reweighting. Both EWN variants yield lower point estimates of aggregate MSE than IN in the matched forecasting experiments, including those with homogeneous heads. Static EWN also improves on IN in the raw-lag control. EWN avoids a discrete search over suffix lengths, but the comparison with head type selected by validation establishes neither an accuracy advantage nor equivalence to trailing-window normalization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.