acceptodds
Under review as a conference paper at ICLR 2027

The Representation Cost of Conditioning

Abstract

Approximate inference in Bayesian neural networks requires choosing which weight dependencies to represent. We study shallow networks with a Gaussian prior, fixed data, Lipschitz activations and smooth convex likelihood energies of at most quadratic growth. Products over whole neurons allow dependence within each neuron; in Gaussian regression, their limiting optimum recovers the posterior mean but retains prior fluctuations. For mixtures of products over individual weights, the minimum log component count at width is to approach a strictly better whole-neuron variational value and to approximate the full weight posterior in KL, at fixed tolerances below the respective positive gaps. The factors may be non-Gaussian. For ReLU, full Gaussian weight distributions with bounded prior KL can contract one direction of limiting uncertainty while missing mean directions available to whole-neuron products. The optimal ELBO ordering changes with the labels. In the width limit, the optimal full Gaussian can give a conditional forecast in the wrong direction even after both predictive marginals are made exact. Finite fits, a covariance intervention that preserves the mean function and a separate preference-head study illustrate consequences for prediction and evidence-based prior selection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.