Residual Geometry Separates State Motion from Predictive Change in Weight-Tied Iterative Refinement
Abstract
For weight-tied networks, the relative hidden-state residual is commonly applied as a convergence diagnostic, yet it can remain large when predictions have frozen and can look similar while predictions are still changing. We explain this mismatch by decomposing a raw update into shared scale, unit-wise scale heterogeneity, and within-unit direction. The decomposition is exact and exposes which motion a readout can see. Across nine 3D refiners, the terminal residual clusters near and consists almost entirely of shared scaling, with no label flips. The fully retained double-precision bundles attribute the late directional plateau in the float32 runs to a storage floor. An extrapolating prefix-sums model and three pretrained LoopFormer checkpoints instead remain orientation- and output-active. Equal-energy interventions confirm that component energy and decision sensitivity are different: direction flips 6–13% of 3D points and 63–75% of LoopFormer tokens, while shared scale flips none. A projective constraint makes the realised state contract but does not stabilise prediction. These results motivate a geometry-and-readout halting gate that requires resolved orientation and output change to settle and refuses unsafe truncation of an unfinished internal clock. On held-out LoopFormer sequences, that refusal is necessary: matched-cost dynamic rules lose more than a nat while the native schedule loses only nats at 25% lower compute.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.