When Does Per-Token Halting Help? Adaptive Token-Wise Damping for Fixed-Point Reasoning Models in Looped Transformers
Abstract
Fixed-Point Reasoning Models (FPRM) use a parameter-free halting rule based on the maximum token residual. A single slow-converging token can therefore keep an entire sequence active. We introduce Adaptive Token-Wise Damping (ATWD), which adapts damping independently at each token and periodically restores every token's initial step-size. This refresh guarantees bounded intervals between full-strength updates while iteration continues; it does not establish dynamical stability or reduce shared-block evaluation cost. Across three tasks, ATWD's benefits depend on the setting. Under causal attention, ATWD improves mean accuracy and observed cross-seed stability on the alternating group , but on the harder symmetric group it fails to reach the baseline's late generalization transition within the matched training budget. Scalar-shaped replays show rapid residual convergence; separate prediction diagnostics suggest low-diversity collapse. On bidirectional Sudoku, ATWD exceeds a hard-freeze-training ablation on sequence accuracy in all three seeds tested (+25.5% relative in the ratio of means), but trails the scalar baseline in mean accuracy and cross-seed stability. Because this ablation changes both training-time decay and refresh, it does not isolate refresh's causal contribution. The hard-freeze-trained models also produce fewer duplicate-constraint violations among their incorrect grids, revealing a trade-off between exact solution accuracy and this measure of failure severity. These results identify task-dependent benefits and limitations of token-wise damping, while distinguishing numerical convergence from solution correctness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.