acceptodds
Under review as a conference paper at ICLR 2027

When to Trust Counterfactuals: Reliable Reinforcement Learning with Partially Learned Causal Models

Abstract

Causal reinforcement learning can improve sample efficiency by reusing factual trajectories to learn from alternative counterfactual outcomes. However, this benefit depends on the reliability of the structural causal model (SCM) used to generate counterfactual trajectories. We study this problem through Counterfactually Guided Proximal Policy Optimization (CG-PPO) with a partially learned SCM, where known mechanisms remain explicit and the unknown mechanism is represented by a model of the other agent (MOA) trained with response-equivalence supervision. We find that CG-PPO can sustain the benefit of counterfactual learning when the SCM is reliable, whereas this benefit diminishes as training proceeds when counterfactual rollouts rely on an imperfect learned SCM. We show that factual response correctness strongly predicts behavioral correctness after counterfactual branching, separating high and low reliability states by 56.4 percentage points after one step and 22.4 points after 40 steps. This discrimination remains positive across the observed range of factual–counterfactual trajectory divergence. We use this signal to introduce reliability-aware branch-state sampling through Consecutive Correctness Weighting (CCW) and Censor-Aware Percentile Weighting (CAPW). When counterfactual PPO updates are interleaved with factual PPO updates, CAPW raises mean win rate over the final six evaluations from 57.0% to 63.0% relative to matched uniform sampling, with higher Final MA6 in all five matched training seeds. It also achieves a 13.7% relative improvement in normalized learning-curve area with respect to factual environment exposure over factual-only PPO, with improvements across all five matched seeds. Under stochastic dynamics, factual conditioning additionally improves learning under the tested SCM misspecification but provides no consistent advantage when the SCM correctly models the exogenous-noise distribution. These results show that locally measured reliability can guide counterfactual training-data allocation when reinforcement learning relies on partially learned causal models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.