acceptodds
Under review as a conference paper at ICLR 2027

Beliefs Beyond Posteriors: The Predictive Cost of ELBO Binding

Abstract

Posterior-seeking training need not produce the best predictive distribution in its family under model misspecification. For a last-layer weight belief, we locate the restriction responsible for this gap and quantify its cost. At the last layer the negative ELBO decomposes exactly into the belief’s predictive log scores, the prior KL, and a binding: a sum of per-observation Jensen gaps that constrains predictive dispersion through posterior covariance. Removing the binding gives the established KL-regularised form of predictive-score training; a parameterisation-dependent factor-normaliser form performs similarly at the last layer. At fixed predictive mean, the leading restricted direction is the variance at each input, and its local value is the classical variance-component score evaluated at the plug-in predictor. This yields both a diagnostic and a price: in the flat-leverage limit, the Gaussian cost is the Jensen gap of a representable residual-variance profile, whereas a likelihood without a dispersion parameter pays the full expressible improvement over the plug-in. Controlled experiments find typically tenths of a nat for Gaussian regression but units to tens of nats for overdispersed Poisson counts. In an exploratory, prospectively specified validation on 28 count sets, 27 unused elsewhere, a screen computed from the point-prediction head alone agreed with the measured outcome on 24 of 25 decided sets (an always-open rule: 18), and its price proxy, once fixed, was within a factor of three of the predictive-score gain of the fitted free profile on all 10 open sets tested afterwards. Score training changes the inferential target: it yields classical per-input overdispersed heads whose covariance represents residual dispersion, not posterior or epistemic uncertainty.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.