When Microbatching Changes the Training Rule: Statistic Scope and Gradient Bias
Abstract
Gradient accumulation is widely used to fit a large training batch into memory and is normally assumed to preserve the update. This equivalence silently relies on splitting the batch without changing the quantities used by the loss. Many modern training rules violate this condition: reward and advantage normalization, contrastive losses, and metric-based objectives compute statistics shared across examples. If those statistics are recomputed inside each microbatch, the compute partition also changes the training rule—even when the global examples, effective batch size, and optimizer step are unchanged. We isolate this effect by varying only the set of examples used to compute the shared statistic, which we call its statistic scope. Two intuitive mechanisms explain the resulting mean-gradient shift: self-inclusion, where an example helps determine the statistic used to evaluate itself, and nonlinear plug-in bias from applying nonlinear functions to estimated statistics. We show that the shift grows predictably as statistic scope narrows and corresponds, to leading order, to optimizing a different effective objective. Across controlled non-decomposable objectives and frozen Qwen3-4B rollouts, the measured effect follows the predicted scaling. On Qwen3-4B, four separately normalized microbatches of 32 produce a 14.74% mean-gradient gap relative to one statistic over all 128 examples, and 20 paired SGD steps produce a path separation equal to 20.4% of the global-statistic displacement. Finally, global statistics can be computed over the full batch while backpropagation remains chunked, preserving the exact global-statistic gradient with chunk-sized activation memory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.