Adaptive Supervision Helps Tiny Recursive Models Learn Better
Abstract
Tiny recursive models (TRMs), as a class of latent depth-recursive models, demonstrate powerful reasoning capabilities and provide promising avenues for scaling reasoning capabilities in limited model size. Despite remarkable progress, training recursive models faces challenges in how supervision is applied across thinking steps. Supervising only the final output provides sparse learning signals, while TRMs address this by applying full supervision at every thinking step. However, at early thinking steps, full supervision can be noisy since the model may need more steps to derive a complete solution. Building on this intuition, we introduce AS-TRM, a simple yet powerful training method for TRMs with Adaptive Supervision. The core idea is to use the outcome of each thinking step to partition the training signals into focus and non-focus sets. We analyze AS-TRM through gradient and loss landscape diagnostics, and show that it reduces supervision noise and leads to faster convergence. Despite its simplicity, AS-TRM delivers substantial gains in both training speed and final performance: on Sudoku-Extreme, it achieves a 7.9-point accuracy gain (11.3% relative) and speeds up training by up to 1.79x; on Maze-Hard, AS-TRM enables flip and rotation augmentation, improving final accuracy by 6.3 points over an already strong baseline; and on ARC-AGI-1, it achieves up to 1.51x speedup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.