Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents
Abstract
LLM agents increasingly retrieve externally curated skills-procedural instructions retrieved at decision time-to improve performance on long-horizon interactive tasks. Existing skill libraries are typically improved in one of two ways: updating model weights-costly, and infeasible for frozen or API-served models-or sharing a single model-agnostic library across backbones with substantially different capacities. However, our controlled experiments across multiple model scales show that skill effectiveness is strongly model-dependent: a skill that benefits one backbone can harm another. Motivated by this, we take a third route-adapting the skills to each model rather than the model to the task-and propose MASA (Model-Aware Skill Alignment), which aligns skills with each target backbone while keeping agent weights frozen. MASA lets skills self-evolve from environmental reward-without human supervision-in two stages: (1) a hierarchical skill evolution pipeline in which a stronger teacher LLM iteratively rewrites general and task-specific skills using hill climbing and UCB-driven tree search, guided by model capability profiles; and (2) a lightweight model-conditioned skill rewriter, trained on these trajectories, that amortizes this search into a single forward pass. Experiments across three interactive environments and four backbones show that MASA consistently achieves the best overall performance, with gains of up to 25.8 points over the strongest baseline. The learned rewriter further generalizes to unseen tasks and environments without additional search, consistently outperforming a much larger teacher LLM at a fraction of the inference cost. Our code is publicly available at https://anonymous.4open.science/r/MASA-BE5F/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.