acceptodds
Under review as a conference paper at ICLR 2027

From Intervention to Skill: Evolving Small–Large Language Model Collaboration

Abstract

Small–large language model collaboration improves inference efficiency by assigning routine decisions to a smaller model and invoking a stronger model when additional capability is needed. Existing approaches primarily optimize this allocation for the current interaction, leaving the value of a large-model intervention largely transient. We study whether past interventions can instead enable collaboration to evolve. We introduce Intervention-to-Skill (I2S), which converts execution-supported large-model interventions into reusable skills. I2S localizes the decision opportunity at which stronger assistance is requested, grounds its effect in environment feedback, and uses trajectory context to identify a reusable procedural residual. The resulting skills persist across tasks and evolve with new evidence, allowing the effective small-model policy to improve while both acting models remain fixed. Across ALFWorld, WebShop, and ScienceWorld with Qwen3-8B paired with Qwen3-32B, DeepSeek-V4-Flash, and Gemini-2.5-Flash, I2S consistently improves task quality over strong routing and skill-based baselines. Compared with SkillGen, task-quality gains reach 27.1 score points and weighted token usage is reduced by up to 63.2%. As skills accumulate, task quality increases while large-model involvement decreases, and acquired skills are repeatedly reused in later small-model decisions. These results establish small–large collaboration as a cumulative adaptation process in which previous assistance reshapes the future division of labor between models. Code is available in the Supplementary Material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.