Learning from Teachers without Training: Functional Geometry for Capability Transfer
Abstract
Post-training endows pretrained models with advanced capabilities but typically relies on iterative rollout generation and gradient-based optimization, making methods such as On-Policy Distillation (OPD) computationally expensive. We investigate whether a Base model can directly acquire a capability already learned by a post-trained Teacher, without reproducing the training process through which that capability was obtained. We find that parameter-space displacement poorly reflects the actual computational consequence of a Teacher update. In contrast, evaluating the same update on Base activations induces a functional geometry that better captures its nonlinear consequence and remains consistent across different activation domains. Motivated by these observations, we propose Functional Geometry Calibration (FGC), a training-free framework that calibrates channel-wise Teacher updates under the Base-induced functional geometry and restores the response magnitude of the corresponding Base channels. FGC requires neither rollout generation nor gradient-based optimization, and independently calibrated capability updates can be directly composed to integrate complementary capabilities from multiple Teachers. Experiments on multimodal reasoning and visual perception show that FGC enables a Base model to acquire capabilities from individual post-trained Teachers and to compose complementary capabilities from multiple Teachers using only checkpoints and a small set of calibration inputs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.