Where to Supervise Algorithm Learning: Same Label Budget, Different Outcomes
Abstract
A label budget specifies how much supervision a learner receives, but not where it can use that supervision effectively. We study this distinction in contextual permutation composition, where the same recurrent attention model can represent an exact algorithm through either hidden-vector or probability-based state feedback. A paired mask intervention retains or removes first-step supervision while giving every example exactly two intermediate labels and an endpoint label. Across eight fresh data blocks and two paired initializations, this substitution changes held-out eight-step accuracy by 87.41 percentage points with probability-based feedback and 55.00 points with hidden-vector feedback at 8,000 updates. Successful full-supervision controls support a 32.41-point difference between these placement effects. Yet training-length success and compositional generalization separate sharply: all 16 probability-based retain models achieve perfect complete-trajectory accuracy on the evaluated 32-step sequences, whereas none of the 48 hidden-vector models reaches 95%. The successful probability-based behavior also survives cutting gradients through the carried state. An earlier 4,000-update comparison and a full-supervision continuation show why training budget changes the interpretation of interface comparisons. Together, these results make supervision placement a measurable design choice and show that label count, state feedback, and optimization budget must be evaluated jointly when studying algorithm learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.