acceptodds
Under review as a conference paper at ICLR 2027

Trajectory-Aligned Gradient Suppression for Test-Time Generalization in Recurrent Transformers

Abstract

Test-time compute scaling has emerged as a powerful paradigm for improving reasoning in large language models. Recurrent transformers offer a natural mechanism for such scaling by repeatedly applying a shared computation block, producing a sequence of latent states that can be unrolled for additional steps at inference. However, additional recurrence often struggles to improve performance beyond the training depth. We study this limitation through a stability–progress trade-off in the latent state dynamics: the model must learn a recurrent operator that remains stable when extrapolated to greater depths, while still making step-by-step progress along the latent trajectory required for multi-step reasoning. We examine this trade-off through the geometry of training gradients. A first-order expansion of the loss under additional recurrence identifies the component of the backpropagated state gradient aligned with the local tangent of the latent-state trajectory as a local channel of depth sensitivity. This trajectory-aligned gradient has a dual role: it can support sequential progress along the recurrent path, but it also determines how the loss changes when computation is extended beyond the training depth. We formulate a curvature-weighted local signal filtering objective for this component and derive a corresponding closed-form suppression rule. Motivated by this analysis, we propose Trajectory-Aligned Gradient Suppression (TAGS), a backward-pass intervention that selectively attenuates tangent-aligned state gradients while preserving transverse gradients for iterative refinement. TAGS adds no parameters, modifies only training-time gradients, and leaves inference unchanged. Across structural reasoning benchmarks (maze navigation, Sudoku, graph shortest path) and symbolic reasoning tasks, TAGS improves test-time compute generalization over backpropagation-through-time (BPTT), truncated BPTT, and fixed-point baselines

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.