Beyond Training Support: Parameterization Selects Fine-Tuning Endpoints
Abstract
Fine-tuning on a narrow support often does not determine a unique model function. Many updates can match all training predictions yet behave differently outside the observed support. Optimization therefore does more than fit the data: it also selects how the model extends beyond the training set. We formalize this choice through a parameterization-conditioned endpoint selector. Specifically, for linear adaptation coordinates, the selected endpoint is an explicit metric-minimum right inverse. For balanced signed-square coordinates, it is a hypentropy-minimum interpolant, with initialization scale controlling a transition between sparse and Euclidean regimes. We then construct parameterizations that produce identical training-prediction trajectories at every step but converge to different endpoints. Next, we derive an analytic boundary for when their held-out ordering reverses. Conversely, we characterize when two selectors agree on a fixed held-out witness and show that, once the training support identifies that witness, all interpolating endpoints must agree. Finally, experiments in solvable systems, frozen representations up to 30B parameters, and matched-fit Transformer adaptations separate endpoint selection from expressivity, convergence, and learning amount. Together, the theory and experiments explain why the same narrow-domain fit can lead to different functions, and when additional support removes that freedom.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.