acceptodds
Under review as a conference paper at ICLR 2027

Falsify Before You Amplify: Conservative Cross-Fitted Credit for Self-Evolving Language Agents

Abstract

Hindsight skills provide dense supervision for language agents, but explaining a completed trajectory does not establish which actions deserve credit. When the same trajectory supplies both a skill and its supporting evidence, reconstruction can be mistaken for guidance. We introduce EvoCred, a reinforcement learning method that separates the authority to establish credit from the ability to allocate it. A frozen policy evaluates each target's fixed actions under a skill generated from another rollout of the same task. Evidence strength and compatibility with verifier-derived advantages govern the skill's influence. Admitted evidence then redistributes token advantages through a positive, bounded transformation that preserves their signs and each trajectory's total absolute credit. The verifier establishes the credit budget; hindsight only reallocates it. Across ALFWorld, Knowledge QA, and WebShop with three Qwen backbones, EvoCred ranks first on ten of twelve aggregate comparisons and outperforms GRPO, SEED, and AgentOPSD on all twelve. On Qwen3-1.7B, it improves ALFWorld macro success from 42.1% to 91.2% and WebShop success from 38.3% to 76.5% over GRPO, without additional environment rollouts or inference-time skill access.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.