Is In-Context Search Worth Its Compute?
Abstract
Training language models on search tasks can be substantially improved by providing supervision via search-augmented (SA) intermediate traces that imitate sequential search. Yet, this requires processing longer sequences than solution-only (SO) training, and prior results leave it unclear when this additional supervision is compute-efficient. We therefore ask: when is learning to imitate search more compute-efficient than learning solutions directly? We study this question on maze navigation and Countdown by systematically varying three factors: training scale, the structure of the underlying search problem, and the structure of the intermediate supervision. We identify two different scaling regimes. On maze navigation, SO supervision reaches low error with less training compute than search imitation; we term this a search-optional regime. On regular Countdown, SA supervision with depth-first search traces is substantially more compute-efficient than SO across the training budgets we evaluate; we term this a search-efficient regime. The advantage of search traces increases on modular Countdown, where the magnitude-based heuristics perform at chance level. Finally, training on incomplete search records reduces the efficiency advantage of SA learning, while semantically unrelated traces behave differently across tasks. These results show that the relative efficiency of SA and SO supervision depends jointly on training scale, problem structure, and supervision quality. Exploiting this insight, we show that selective SA supervision achieves low error with less training compute in joint training on both tasks than uniformly applying either SA or SO supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.