acceptodds
Under review as a conference paper at ICLR 2027

Trainable Affine Parameters in Normalization Can Silently Harm Analog Training

Abstract

Analog in-memory computing (AIMC) offers a route to energy-efficient training by accelerating matrix–vector multiplications (MVMs) in memory. With finite-resolution analog MVMs in both forward and backward passes, trainable normalization gain and bias before the final analog linear layer can degrade training, despite remaining digital and unquantized. Part of the loss gap relative to fixed gain and bias persists when the trained head is replayed with exact multiplication, indicating damage retained after training. To understand this failure, we show how affine parameters can increase absolute logit error through changes in input scale and direction, despite dynamic max-absolute scaling. Training can fail to correct these effects because the supplied gain gradient can predict a loss response opposite to that measured by forward evaluations. These findings motivate parameter-free normalization before the digital-to-analog converter, with any needed affine placed after the analog-to-digital converter. Normalizing the digitized output can also mitigate the damage. Language and vision experiments reproduce the failure and its repairs across converter precisions, optimizers, and model scales, with repair benefits depending on the task, optimizer, and operating point.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.