Rethinking Agent Skill Orchestration: Do we Need to Compile Skills Before Each Execution?
Abstract
As agent skills become widely adopted, the prevailing practice is to retrieve relevant skills for each task and orchestrate them on the fly. When an agent repeatedly solves tasks of the same type, however, recompilng skills for every instance incurs extra LLM calls and latency, and discards the memory accumulated in earlier executions. Prior work compiles skills into dependency graphs or induces reusable skills from trajectories offline, yet still leaves orchestration to test time, performed anew for each task. We observe that successful executions of the same task type follow a highly consistent composition, so composition can be done once, at induction time. Building on this observation, we propose FlatSkill, which compiles an agent's skill-guided trajectories into one prefab per task type, pairing a slotted, task-level procedure with rules verified against the trajectories. At test time, the agent only selects the matching prefab and fills in its slots. On ALFWorld, ScienceWorld, WebShop and MedCalc-Bench, FlatSkill outperforms existing methods in all evaluation settings with DeepSeek-V3.2, while taking fewer interaction steps than the strongest baseline. Prefabs compiled by a stronger model can also be reused directly by a smaller one, demonstrating the benefit of amortizing skill composition across executions rather than repeating it for every task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.