Cached Boundary Transport for Efficient Activity Inference in Deep Predictive Coding
Abstract
Predictive coding (PC) networks infer their hidden activities before every weight update, and in deep networks this inner optimization becomes the bottleneck of training. We show that the difficulty is geometric rather than iterative: the activity landscape contains low-curvature directions that span the whole layer chain, and a roughly 28-fold larger gradient budget leaves the test accuracy of 256-layer networks unchanged. A residual PC chain, however, supplies its own solver for these directions. Linearizing the predictions at the feedforward state yields a batch-shared Gauss–Newton preconditioner whose inverse is applied exactly by one backward and one forward recurrence over the layers, plus a small cached output-space correction. Cached Boundary Transport (CBT) uses this solve for the chain-wide component and a short gradient-descent refinement for the remaining local error; an energy acceptance criterion and a restricted Gauss–Newton fallback handle nonlinear mismatch. From 16 to 256 layers, CBT keeps inference cost nearly constant in gradient evaluations while first-order and Newton-type solvers grow with depth. On trained networks it brings the parameter gradient substantially closer to that of an exact second-order oracle, and in end-to-end training it is more accurate than depth-proportional inference at every tested depth with up to fewer FLOPs. These results point to chain geometry, not the step budget, as the bottleneck of deep residual PC inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.