acceptodds
Under review as a conference paper at ICLR 2027

Learned Straight-Through Estimators for Binarized Neural Networks

Abstract

Binarized neural networks (BNNs) training requires an approximation of the gradient of the sign function, whose derivative is zero almost everywhere. Current methods replace the sign gradient with some hand-crafted surrogate, such as the straight-through estimator (STE) and its variants. The STE shape is decided in advance, whether held constant or varied by a progressive schedule; this is despite the distributions it acts change over layers and epochs. To address this, we propose the Learned STE (LSTE), which lets the training objective tune the shape of the STE in every quantizer. Specifically, we give each weight quantizer and each activation quantizer one learnable parameter that controls the shape of their STE. These parameters act only in the backward pass, so the current mini-batch gives them no gradient; we train them instead on the loss of the next mini-batch, which depends on the weight update they shape. The forward pass is unchanged, so LSTE combines with different forward quantizers and adds no cost at inference. The same mechanism also selects between published estimators by learning the ratio that mixes them, and shows that the preferred one depends on the quantizer. Experiments on CIFAR-10, ImageNet, and 1-bit language models of the Llama family trained on C4, show that LSTE improves on state-of-the-art methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.