Function Coordinates are Predictive
Abstract
Measuring a model's predictions over data offers a lens into understanding the functions it learns. Recent work places models in a shared coordinate system based on their predictions on a fixed dataset, providing a lens into the functions they learn. Yet, existing systems fall short in two ways. Some position models by their similarity to a specified set of reference models, describing each new model only through the information captured by the set, which can misorder its relationships to other models. And all existing systems are temperature-sensitive: two models with identical decisions but different temperatures can lie farther apart than two models that disagree with the same temperature. We propose a simple coordinate system that addresses both shortcomings by construction: models are placed by their own predictions alone, with temperature removed. Across 1,000+ full training runs spanning 15.6K checkpoints probed on ImageNet-1K, we ask what a checkpoint's position reveals about its training beyond how well it performs: dataset and patch size are linearly decodable at any training stage, model size only weakly, and optimizer settings only through accuracy and loss. Moreover, trajectories are predictive: a non-parametric estimator recovers a run's final and per-class accuracies to within a few points from its first 10-25% of training, while its per-input class predictions and confusion matrix become predictable in the second half of training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.