As a Hidden Hyperparameter, Hardware Changes the Training Algorithm
Abstract
Splitting a training update across GPUs or microbatches should save time and memory, not change the algorithm. Yet in practice, it can. We introduce Partition Invariance: repartitioning the same data must preserve the training update in exact arithmetic when the specified algorithm is unchanged. We prove a sufficient condition for this property and use its premises to guide diagnosis. PartitionCheck tests this property by comparing the same update across two partition configurations, allowing for floating-point differences. It also checks each execution against an independent reference implementation, since matching executions can share the same error. Beyond detection, we characterize how partitioning changes the algorithm. Under stated conditions, constant gradient scaling changes the effective learning rate under SGD but weight decay under LARS. In five codebases, we test whether code-derived rewrites reproduce the affected runs. Across 31 frameworks and research codebases, including Transformers, DeepSpeed, and veRL, we identify 43 implementation errors; six affect every partition alike. These errors can undermine research conclusions even without degrading model quality. In SimCSE, two GPUs double the relative weight of an auxiliary loss. Fixed padding bypasses it, making an ablation compare identical training runs. In a preregistered experiment with a 134M-parameter Mamba-2 model, the affected tensor-parallel implementation yields 0.189 nats higher test loss than its repaired counterpart when both are evaluated under the declared architecture. Reporting the partition is not bookkeeping; it is part of reporting the method.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.