From Advantage Perturbations to Counterfactual PPO Guidance
Abstract
Proximal Policy Optimization (PPO) relies on advantage estimates to reinforce useful actions, but in sparse-reward tasks these estimates can be delayed and noisy, making it difficult to identify which earlier actions led to success. An external oracle, such as a planner, expert, safety checker, or simulator, can provide action-level information not available from the environment reward. We study advantage perturbations as an interface for adding oracle information directly to PPO’s policy update, without changing the environment reward or executed actions. We show that state-only perturbations do not shift the expected local policy gradient at the rollout policy, whereas action-relative perturbations can change preferences between actions at the same state. Controlled CartPole experiments confirm this mechanism and show that reversing the perturbation reverses the learned preference for a corrective action. We then use task-informed manual oracles to guide PPO on four sparse-reward MiniGrid tasks. Since manually specifying preferred actions requires task-specific knowledge, we next introduce a more general oracle based on local counterfactual action evaluation, which compares the outcomes of alternative actions from selected rollout states. Beyond adding this counterfactual signal at the advantage level, we also test it in the TD update and directly in the policy-loss objective. These interventions show task-dependent effects relative to standard PPO, with improvements in several task–intervention combinations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.