Different Bonuses, Similar Outcomes: Understanding Auxiliary Reward Non-Separation in PPO
Abstract
Auxiliary-reward methods are usually distinguished by the quantity that generates their bonus, yet the bonus itself is not the object optimized by a policy-gradient method. We study a general learning question: how can mathematically different auxiliary objectives produce nearly indistinguishable behaviour after integration into the same policy-optimization pipeline? Six PPO-based systems are compared in a controlled dynamic-survival task: PPO alone, intrinsic curiosity (ICM), random-network distillation (RND), hashed pseudocount exploration, advisory planning, and ICM with advisory planning. Across 40 matched seeds, 240 policies are trained in 120,000 episodes and evaluated in 336,000 held-out episodes. Raw intrinsic signals remain clearly distinguishable; their mean per-step magnitudes span more than three orders of magnitude, while stochastic test survival occupies only a 32.63% to 33.25% range. The hybrid is numerically first, but its paired effects are only to percentage points; every Holm-adjusted survival test has , and every confidence interval excludes the prespecified five-point benefit. The ICM-by-planner interaction is also unsupported. We formalize an effective-update view in which generalized advantage estimation, episodewise standardization, PPO clipping, and mapping through policy score vectors form a many-to-one map from raw auxiliary rewards to actor updates. The experiment does not establish equal gradients or general algorithmic equivalence. It does establish a controlled case in which semantic and numerical differences among auxiliary signals fail to yield practically distinguishable downstream outcomes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.