Nearly Opposite Gradients in Shrinking Regions of Looped Networks
Abstract
Looped networks increase computational depth by repeatedly applying shared modules, whereas truncated backpropagation differentiates through only a subset of these applications. We study how this mismatch affects gradient direction in a normalized hierarchical loop with low-level iterations per cycle and truncation of the first low-level group. Increasing depth can make the truncated gradient nearly opposite to the full gradient, even as the reversal region contracts. For a fixed task and neighborhood in the parameter space, the ascent region has thickness and initialization probability . A signed expansion shows that omitted parameter contributions accumulate at scale and compete with the local full gradient. A correction based on data moments and current parameters reduces the remaining thickness to . We also characterize when discrepancies between cycle contributions persist after low-state convergence. Solver relaxation preserves the limiting alignment, whereas suitable block weights can eliminate an ascent interval. We verify finite-depth boundary predictions in a nonlinear residual family. At trained looped-Transformer checkpoints, restoring early contributions achieves higher cosine similarity with the full gradient than any nonnegative block rescaling and produces larger cross-entropy reductions under matched updates. Thus, greater loop depth can reduce the prevalence of gradient reversal without reducing its severity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.