Langevin Temporal Predictive Coding: From Point Estimates to Posterior Distributions
Abstract
Temporal predictive coding (tPC) is a model of sequential memory with local inference and Hebbian learning. Its inference, however, is deterministic. At each time step it propagates a point estimate of the latent state rather than a posterior distribution. We add Langevin noise whose scale is set by the fluctuation–dissipation theorem (FDT). The resulting Langevin tPC samples the conditional posterior at each time step. A Langevin step has two hyperparameters, the preconditioner and the step size . We compare two preconditioners, the identity matrix and the transition covariance , and relate them through the effective step size. Because latent directions relax at different rates, we also propose a per-direction step size that shrinks as curvature grows. On synthetic targets, the energy and the KL divergence decrease as the FDT predicts. Under anisotropy, the per-direction step size reaches lower KL divergence than identity preconditioning. Langevin tPC also represents multimodal posteriors, which MAP inference collapses to a single mode. We then evaluate our method on twelve regression datasets with Gaussian-mixture priors, where the exact posteriors are known and ten of them are multimodal. Across all twelve datasets, Langevin tPC approximates the exact posterior more closely than an ensemble of MAP estimates. On five video benchmarks, where posteriors are unimodal, it matches standard tPC. Adding noise thus extends tPC from point estimates to posterior distributions, and from sequential memory to probabilistic inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.