Best-of-Depth: Learning Exploratory Dynamics in Looped Transformer
Abstract
Looped Transformers increase inference-time computation through repeated latent-state updates with shared parameters. We propose training these updates to increase the utility of candidate solutions along a single trajectory. Our method, Best-of-Depth (BoD), minimizes the loss of the best prediction across depths, encouraging a match to the target anywhere within the depth budget without architectural changes. Experiments on N-Queens completion and graph coloring show improved solution discovery under greedy decoding relative to training that supervises only the prediction at the maximum recurrent depth. On N-Queens, new solutions continue to emerge beyond the training depth before discovery saturates. Code-generation experiments also show higher success across evaluated depths than maximum-depth training. Fixed recurrent-step-budget comparisons favor combining depth with multiple trajectories. These findings show how training a recurrent trajectory as a candidate set enables depth-wise solution discovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.