acceptodds
Under review as a conference paper at ICLR 2027

Learning What Matters: Counterfactual Policy Learning for LLM Agents

Abstract

Large language model (LLM) agents increasingly improve through accumulated interaction experience, yet it remains unclear which decisions within a trajectory are worth learning and reusing. Successful trajectories may contain many incidental actions, so trajectory-level memorization cannot reliably distinguish essential decisions from irrelevant details. We propose Counterfactual Policy Learning (CPL), a decision-level experience learning framework that separates candidate identification from policy validation. CPL contrasts successful and failed trajectories to identify candidate decision differences, treating this observational signal only as relevance rather than causal evidence. It then abstracts candidates into intervention-testable policies and performs same-task interventions that replace policy-consistent decisions with alternative actions. Only policies with positive intervention effects and sufficient admission confidence are admitted into policy memory; unseen-task performance is used solely for downstream evaluation. On ALFWorld, policy-guided interventions recover 10.8% of failed episodes versus 0% for random replacements; on LoCoMo, CPL achieves 32.65% overall Semantic Similarity. Ablations show that type-matched retrieval acts as a firewall against negative transfer, reducing ReusePerf from 53.86% to 43.81% when replaced by random retrieval. CPL thus reframes experience learning from “what happened?” to “which decision is worth learning from?”Our code is available at https://github.com/linda-2-hh/Counterfactual-Policy-Learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.