acceptodds
Under review as a conference paper at ICLR 2027

One Skill, Many Models: Learning Representation Compatibility from Sparse Behavioral Tests

Abstract

Agent skills can transfer poorly across large language models (LLMs) even when their procedures are unchanged: models may respond differently to semantically equivalent choices in instruction structure and executable organization. We introduce SkillProbe, a budgeted framework that diagnoses these model-specific representation preferences before skill adaptation.SkillProbeBank provides 156 matched probes and 360 semantically equivalent variants spanning 14 instruction and executable factors. SkillProbe uses a hierarchical Bayesian profile to track preference effects and uncertainty, then selects probes according to their expected improvement to representation decisions per unit cost. The resulting profile can condition existing methods for skill migration, generation, cold start, and version-drift repair. At 20% of the full-bank budget, SkillProbe achieves 84.3% agreement with full-bank representation decisions using 198K tokens, outperforming adaptive low-budget baselines at lower cost. Integrated with strong skill learners, it achieves the best average rank across ALFWorld, WebShop, SkillsBench, and SWE-Skills-Bench. Across seven additional transfer settings involving Gemma, Llama, Qwen, and a GPT-4o target, SkillGraph+SkillProbe outperforms SkillGraph in 27 of 28 paired benchmark comparisons. These results suggest that sparse behavioral diagnosis provides a reusable signal for adapting skills across model deployments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.