Hinted Scaling Laws
Abstract
Downstream scaling laws let developers predict benchmark performance at large scale from small, cheap proxies. This prediction fails when the small proxies score zero because their accuracies give no signal to extrapolate to larger scale. We introduce hinted scaling laws, which use privileged information to turn a benchmark into a continuous sequence of easier versions so that even the small proxies score above zero. We find that accuracy rises smoothly and predictably with the hint level, enabling new types of scaling laws that operate below the noise threshold. Within pretrained and post-trained model families, hinted scaling laws fit with the smallest models reduce the median forecasting error for larger models from 21 percentage points for a standard scaling law to 6. Hinting also improves the sample efficiency of model ranking. It widens the gaps in accuracy between models while approximately preserving their order, so a single hinted rollout per problem ranks near-indistinguishable models as accurately as 5 unhinted rollouts. Our results suggest that hinting is a general primitive for improving the accuracy and cost of downstream scaling laws and other prediction tasks in model development.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.