Counterfactual Supervision for Credit Assignment in Language Model Agents
Abstract
We develop a shared turn-credit selector using counterfactual supervision from 4B and 9B policies before and after reinforcement learning. Successful coding trajectories can contain harmful actions, while failed ones can contain useful decisions. To distinguish these local roles, we restore the state before each action and repeatedly execute both the observed action and same-policy alternatives in the environment. Their difference in average correctness–efficiency utility defines a gold credit label. These labels guide development of a parameter-frozen retrospective selector, which supplies the actor's complete turn-level advantage without online counterfactual branching. Each score is shared by the turn's action tokens in a clipped policy objective; training uses no learned value critic, GAE, or additional outcome advantage. Across four evaluation sets and both actor scales, our method exceeds all four trained baselines. SWE-bench Pro gains over PPO are 4.52 and 4.79 percentage points at 4B and 9B, respectively. At 9B, it matches the Base model on LiveCodeBench while exceeding the other trained methods. Credit-removal experiments distinguish effective learning from simple trajectory compression. Matched-policy probes further support improved local action selection under a shared continuation policy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.