acceptodds
Under review as a conference paper at ICLR 2027

Stout: Structured Order-Aware Quantization Training for Large Language Models

Abstract

Layerwise quantization-aware training (QAT) reduces the training cost of large language models by optimizing one layer at a time. However, a fixed input-to-output order assigns adaptation by depth, without considering which layers can best recover from errors introduced by other layers. It can therefore leave layers with limited adaptation ability to handle errors after more capable layers have been frozen. We introduce Oats, Order-aware Adaptation Score, which estimates a layer’s potential recovery from other layers’ quantization errors. Our local analysis explains this recovery, and experiments show that Oats predicts it more effectively than own-layer sensitivity. Based on Oats, we propose Stout, Structured Order-Aware Quantization Training, a two-stage framework for efficient QAT. The first stage computes Oats through shared measurements and trains layers with the model’s output loss in ascending score order, preserving stronger adaptation opportunities for later. The second stage trains scales with fixed quantized weights using ScaleKernel, a fused backward kernel that avoids dense weight-gradient intermediates in global memory. On Qwen3.6-35B-A3B at 2bit, Stout improves LiveCodeBench pass@1 and the BFCL aggregate score by 0.87% and 2.68% over the development-selected ordering baseline. Compared with a matched compiled or fused baseline, ScaleKernel lowers second-stage training time and peak alloated GPU memory by 0.85% and 14.47%, respectively. Our code is available at https://anonymous.4open.science/status/stout-613D

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.