The Truncation Trap: Silent Failure of Fixed-Window Diffusion Likelihood and Information Estimators under Representation Drift
Abstract
Diffusion-based estimators of likelihood and mutual information evaluate an integral of denoising error over signal-to-noise ratio; in practice the integral runs over a fixed, truncated log-SNR window whose placement and data-scale assumptions are frozen at configuration time. We document a latent failure mode of this family: when the representation's scale drifts, most commonly because a jointly learned embedding grows under cross-entropy training, integrand mass silently migrates out of the window and the estimator retains an arbitrarily small fraction of the full-line integral, with no warning. The mechanism is predictive: the integrand peak sits at for representation scale . We demonstrate it end to end on a 306M-parameter diffusion language model whose truncated readout reports 0.78 BPC against a converged dense-grid value of 4.27. Auditing the released code of four published systems, we find that each carries an undocumented guard: per-forward embedding normalization (Plaid), default input standardization (MINDE, MMG), and data-adaptive windows (ITD). Disabling MINDE's guard shifts a scale-invariant MI estimate by 4.2 nats; ablating Plaid's guard in a from-scratch retrain leaves the model intact but sends the fixed-window readout to 278 BPC, whereas a guarded positive control reproduces Plaid's reported 1.12 BPC in three independent ways. We distill a defense kit, consisting of an entropy-floor assertion, convergence-verified windows with measured tail bounds, and a rescale stress test, that detects or precludes this failure mode before a number is reported.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.