Local-Window Contrastive Forward-Forward Learning with Masked Reconstruction
Abstract
End-to-end Vision Transformer training retains a backward graph across the whole network and can become memory-prohibitive on resource-constrained hardware. Restricting backpropagation to a local part of the network reduces this burden, but contrastive Forward–Forward (CFF) mainly supervises pooled representations rather than patch-level visual structure. We introduce contrastive Forward–Forward with masked autoencoding (CFFMAE), which augments local CFF training with masked reconstruction without restoring a full-depth backward graph. At each stage, CFFMAE optimizes a small local window of encoder blocks together with a lightweight decoder, while preceding blocks are evaluated without autograd. We consider layer-wise one-block and pair-wise two-block windows. Unmasked views provide the CFF objective, and masked views provide patch-level reconstruction to the same active block. We separate the encoder input mask from the reconstruction-loss mask, transfer the decoder checkpoint across successive local windows, and backpropagate the CFF and MAE branches separately to limit temporary memory. We evaluate CFF+MAE on CIFAR-10, ImageNet-1K, and the STB hand-pose benchmark, with additional Jetson profiling against CFF-only, full backpropagation, and activation-checkpointed BP. Masked reconstruction consistently improves CFF-only training in the controlled classification experiments. On Jetson, CFF+MAE substantially reduces peak memory relative to uncheckpointed full backpropagation and requires lower per-batch latency and energy than checkpointed BP at comparable peak-memory footprints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.