acceptodds
Under review as a conference paper at ICLR 2027

Why MCPC generalizes

Abstract

Monte Carlo predictive coding (MCPC) infers latent causes by Langevin sampling, which makes it a generative model. We ask when an MCPC network generalizes during training and when it begins to memorize. Its local Hebbian updates average over the sampled cloud of latent states, so the cloud's spread acts as a ridge. For a linear–Gaussian coder we derive this ridge in closed form. It sets a threshold on the data spectrum, and directions farther above it are learned sooner. Population directions lie far above and set the generalization onset. Directions that exist only because the training set is finite lie just above and set a later memorization onset. The network generalizes without memorizing in the window between the two onsets. More data widens this window, and past a critical data size memorization never begins. Longer or more chains move the memorization onset but cannot prevent it. A larger step size can prevent it, but only by also losing the signal. The theory's learning flow reproduces both measured onsets. A pre-registered causal test shows that the ridge, not the sampling noise, sets the memorization onset. On CelebA faces more data likewise delays memorization while sample quality peaks at about the same step. To make MCPC generalize rather than memorize, stop training between the two onsets or train on more data than the critical size.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.