Stout: Structured Order-Aware Quantization Training for Large Language Models
Abstract
Layerwise quantization-aware training (QAT) reduces the training cost of large language models by optimizing one layer at a time. However, a fixed input-to-output order assigns adaptation by depth, without considering which layers can best recover from errors introduced by other layers. It can therefore leave layers with limited adaptation ability to handle errors after more capable layers have been frozen. We introduce Oats, Order-aware Adaptation Score, which estimates a layer’s potential recovery from other layers’ quantization errors. Our local analysis explains this recovery, and experiments show that Oats predicts it more effectively than own-layer sensitivity. Based on Oats, we propose Stout, Structured Order-Aware Quantization Training, a two-stage framework for efficient QAT. The first stage computes Oats through shared measurements and trains layers with the model’s output loss in ascending score order, preserving stronger adaptation opportunities for later. The second stage trains scales with fixed quantized weights using ScaleKernel, a fused backward kernel that avoids dense weight-gradient intermediates in global memory. On Qwen3.6-35B-A3B at 2bit, Stout improves LiveCodeBench pass@1 and the BFCL aggregate score by 0.87% and 2.68% over the development-selected ordering baseline. Compared with a matched compiled or fused baseline, ScaleKernel lowers second-stage training time and peak alloated GPU memory by 0.85% and 14.47%, respectively. Our code is available at https://anonymous.4open.science/status/stout-613D
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.