acceptodds
Under review as a conference paper at ICLR 2027

Curvature-Corrected Activation Gradients for Binarized Neural Networks

Abstract

Binarized Neural Network (BNNs) training typically relies on the straight-through estimator (STE) as a practical approximation of the derivative of the quantization layer applied to the model weights and activations. We identify that the activation gradients in standard STE are biased and that this bias can be decomposed into two components: one related directly to the STE and another related to applying the classical chain rule through a discontinuous quantizer, which propagates the tangent of the downstream loss at the selected quantization level. In contrast, the distributional derivative of the population loss is governed by the finite loss jump between adjacent levels. A second-order expansion of this jump yields a gradient-correction term involving the quantization residual and the diagonal curvature of the downstream loss. We turn this result into a practical backward rule for activation quantizers by approximating the curvature component with an empirical-Fisher proxy and apply the proposed correction during the final 10% of training. We evaluate our method on binarized Llama-style language models trained with recent state-of-the-art methods, including QuEST and CAGE, and observe a 1.0–3.1% perplexity reduction across model sizes up to 1.6B parameters.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.