Dense Supervision for Iterative Neural Computation
Abstract
Iterative Neural Computations (INCs), such as Universal Transformers and Deep Thinking networks, have demonstrated strong capabilities across algorithmic, reasoning, and perceptual tasks, by iteratively applying a shared function multiple steps. Despite their promise, training these systems remains challenging. These models are typically trained by applying a loss only at the final iterate, a recipe we call endpoint supervision. We hypothesize that this recipe is the cause behind two coupled failure modes: (1) gradients from the final loss must backpropagate through the entire chain of repeated updates, producing a learning signal that is both noisy in magnitude and unreliable in direction, and (2) intermediate iterations are left unconstrained, allowing the model to learn a fixed-horizon strategy that breaks down when more iterations are run at test time. We introduce Dense Intermediate Consistency for Endpoints (DICE), an architecture-agnostic framework that addresses both issues by supervising the full trajectory of computation. DICE attaches a shared readout head to every iterate and applies auxiliary losses at each step, providing gradient signals that bypass the deep composition while enforcing that every intermediate state remains semantically meaningful. Across three representative architectures, DICE yields consistent gains, including near-perfect algorithmic extrapolation on prefix sums, pp exact-match accuracy on maze solving, and a inference-time speedup via adaptive halting. Overall, these results indicate that dense supervision is a simple and broadly effective remedy for the failure modes of endpoint-supervised INCs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.