Understanding Agent Skills: How Well Do They Transfer to Smaller Models?
Abstract
Skill evolution gives agents a way to improve over time by turning task-execution experience into reusable procedural guidance. Expressed in text, these lessons can be shared across models without transferring weights. We investigate whether skills developed by frontier models can help much smaller agents generalize to unseen problems while reducing reliance on expensive models. Though text makes this transfer more convenient, smaller models may interpret and follow the same instructions differently from their authors. Across distinct skill-generation frameworks and model families, we find that strong-model skills do not consistently improve smaller executors, and feedback-driven evolution does not consistently outperform direct generation. Analyses of instruction content, model computation, and reasoning trajectories show that responding to a skill does not ensure a better solution. Experiments that introduce guidance at different stages further show that its usefulness can change as a solution develops. Finally, because smaller models are comparatively inexpensive to fine-tune, we compare external skill transfer with internalizing experience through reinforcement learning and on-policy distillation. Training improves performance even without external guidance. After skill-guided training, retaining the document at inference provides further gains, although adding it to a model trained without that guidance can sometimes hurt. Together, these findings show both the promise and the difficulty of transferring experience through skills.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.