Loss Units Change the Algorithm: Consistent Multi-Optimizer AMP
Abstract
Replacing a batch-mean loss with a batch sum, with a compensating learning-rate change, preserves the real-arithmetic SGD update. Under automatic mixed precision (AMP), however, the larger loss can cause one optimizer to skip while another steps. We study how this numerical decision changes multi-optimizer training. Transactional Covariant AMP (TC-AMP) adjusts each loss scale inversely to its loss units and uses joint replay: if any block overflows, no optimizer steps and the same minibatch is retried. We prove exact execution equivalence for disjoint SGD and classical-momentum blocks under power-of-two unit changes. The proof assumes matched scaled backward expressions and normal-range arithmetic. In partial DINOv2 fine-tuning on CUB-200-2011, native AMP produces eight partial updates per sum-unit run and different mean–sum model endpoints. TC-AMP matches parameters, predictions, update events, and reference-unit scaler histories between loss units in each of three seeds. It reaches development validation accuracy, versus for FP32, at its committed-image throughput in this cached-feature workload. A per-block replay baseline also matches model endpoints, but requires eight additional retries. Checkpoint interventions show why this distinction matters: saved scaler state alone can change retry cost even when model and optimizer states match. Loss units can thus change either the updates performed or the computation required to perform them. Consistency tests for mixed-precision training should track update decisions and consumed minibatches alongside final weights.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.