acceptodds
Under review as a conference paper at ICLR 2027

Bad Outcomes Need Not Mean Bad Actions: Counterfactually Verified Pain Memory for Policy Optimization

Abstract

Bad outcomes do not always mean bad actions. Policy-gradient methods update policies using advantage estimates and TD errors, while failure-memory approaches often retain experiences based on observed failures or their severity. When exogenous randomness dominates the effects of individual actions, these signals can mistake bad luck for poor decisions, while later policy updates can erase useful lessons from rare but consequential events. We propose , a two-channel framework built on Proximal Policy Optimization (PPO). The Advantage channel learns from current rollouts, while the Pain channel retains only corrections verified through counterfactual replay, without modifying rewards or advantage estimates. For a candidate failure, we intervene on the executed action and replay the episode in closed loop under the same recorded conditions. A correction enters memory only if the alternative action improves the outcome. Its influence grows when the same corrective direction recurs across distinct realizations. Each stored lesson defines a local, revisable, one-sided retention constraint whose gradient is combined with the PPO gradient only when needed to resist forgetting. We evaluate PAIN-PPO in both controlled simulations and real-world financial markets. In simulation, PAIN-PPO distinguishes actionable policy errors from adverse outcomes driven by exogenous randomness, and its gains largely disappear when counterfactual verification is removed, while selective admission further filters corrections that do not recur consistently. On held-out financial market data, PAIN-PPO demonstrates promising lower-tail risk performance under the prescribed mean-cost budget. Together, these results show that counterfactually verified memory can improve tail-sensitive policy learning by retaining rare but actionable corrections without conflating unfavorable outcomes with poor decisions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.