CHARACTERIZING REPRODUCIBILITY: RUN-TO-RUN NUMERICAL VARIATION IN A NEAR PRODUCTION EXASCALE COMPUTER
Abstract
Reproducible training results require reproducible numerical accuracy across nodes and across time. It is useful in that context for the run-to-run variation of accuracy to be known and reasonably constrained. We characterize the empirical distribution of Top-1 accuracy for ResNet-50 trained on ImageNet-1k. Across 340 instances drawn from 7 cluster wide runs, each using 2052 tiles (171 nodes), final top-1 accuracy spans 0.7062 to 0.7307, a spread of 2.45 percentage points, with relative standard deviation of 0.0056. We show that this variation is not a consequence of scale alone: we show that variations exist even at single rank on a separate A100-based platform with all framework determinism controls enabled that were available at the time of experiment. As a positive control, we report a pipeline change that moves image augmentation to the GPU that costs 3.6 to 8.3 percentage points of accuracy, an effect several times larger than the run-to-run spread. We conclude that single-run accuracy measurements are statistically unsound and recommend adopting a variance-aware range when interpreting large-scale training results. This work also provides practical guidance for reasoning about reproducibility concerns in modern AI workloads like LLM, VLM, etc.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.