A Trail from Low-Precision Sampling: Diagnosing Super-Diffusive FP32–BF16 Drift in Diffusion Models
Abstract
Low-precision inference is now the default deployment path for diffusion image generators: BF16 is standard and FP8 is spreading fast. The rounding error this introduces is usually treated as small, zero-mean noise that averages out over the tens of denoising steps of a sampler. We show it does not average out. Running SDXL from an identical seed in FP32 and BF16, we find that the same-seed latent trajectories do not wander apart like independent noise: their per-step L2 divergence grows super-diffusively, with a log–log growth exponent of 1.04 (diffusive random-walk null = 0.5) and a locatable takeoff around step 25 of 50. This is a robust same-seed phenomenon (n=64 pairs), and we pin down its cause with an unusually thorough elimination: five controls that rule out the tempting statistical explanations one at a time. It is not a mean bias (sign-bias statistic |b| ≈ 0.012). It is not the error’s step-to-step directional correlation: although consecutive error vectors are nearly collinear (cos(e_t, e_t+1) = 0.971), that merely rides the trajectory’s own velocity autocorrelation (0.979), and a magnitude-matched but uncorrelated Gaussian perturbation injected afresh at every step reproduces the super-diffusion (exponent 1.01 ≈ 1.04). And it is not a fixed-magnitude dynamical amplification: the same continual injection at constant per-step magnitude is merely diffusive (0.43), while a one-time kick is damped (0.19). What remains is a growing-magnitude positive feedback: the per-step rounding error scales with the accumulating separation, so as the same-seed trajectories diverge the error grows and drives them further apart—faster than √t, independent of the error’s mean or correlation. Crucially, we do not stop at inferring this feedback but measure it directly: regressing the per-step divergence increment on the current divergence is decisively positive (Spearman 0.89, positive for all 64 of 64 prompts). We package these measurements as DriftProbe, a training-free diagnostic—usable on any UNet or DiT block—that reports the per-step sign bias, the same-seed divergence curve and its onset t*, and the injection controls that isolate the mechanism. We treat the drift as a lens, not a defect to correct, and lay out how t* could serve as an early signal for downstream generation defects. Results are measured on SDXL/BF16 on a single 48 GB GPU (n=64 same-seed pairs; n=16 per injection control); the FP8 arm, the FLUX backbone, and the defect-prediction evaluation are scoped as future work.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.