acceptodds
Under review as a conference paper at ICLR 2027

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation

Abstract

Skills can improve agentic reinforcement learning by reusing past experience as textual guidance, but retrieved skills are not always reliable: they may help in one state while misleading in another. This makes the common privileged-teacher assumption fragile, namely that a skill-conditioned prompt can be treated as a fixed teacher for the no-skill prompt. We propose UCOB, a framework that first learns when to utilize agentic skills via credit-aware on-policy bidirectional self-distillation, then evolves them through utility-aware skill updates and reflection self-training. UCOB constructs two on-policy context views from skill-conditioned and no-skill prompts, compares return-to-go among rollouts sharing the same task and anchor state, and uses the higher-return view as the local teacher. The resulting local credit signal internalizes useful skill-conditioned behavior, corrects misleading skill usage, and further updates the dual-granularity skill memory, informs utility-aware retrieval, and supports reflection self-training. Experiments on ALFWorld, WebShop, and Search-QA show that UCOB outperforms skill-free RL baselines, skill-augmented methods, and self-distillation baselines across model scales, achieving up to 23.5 and 18.0 absolute success-rate gains over SOTA baselines on ALFWorld and WebShop, respectively. Ablations and analyses further validate its core mechanisms, continual adaptation across environments, and modest training overhead.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.