acceptodds
Under review as a conference paper at ICLR 2027

Different Bonuses, Similar Outcomes: Understanding Auxiliary Reward Non-Separation in PPO

Abstract

Auxiliary-reward methods are usually distinguished by the quantity that generates their bonus, yet the bonus itself is not the object optimized by a policy-gradient method. We study a general learning question: how can mathematically different auxiliary objectives produce nearly indistinguishable behaviour after integration into the same policy-optimization pipeline? Six PPO-based systems are compared in a controlled dynamic-survival task: PPO alone, intrinsic curiosity (ICM), random-network distillation (RND), hashed pseudocount exploration, advisory planning, and ICM with advisory planning. Across 40 matched seeds, 240 policies are trained in 120,000 episodes and evaluated in 336,000 held-out episodes. Raw intrinsic signals remain clearly distinguishable; their mean per-step magnitudes span more than three orders of magnitude, while stochastic test survival occupies only a 32.63% to 33.25% range. The hybrid is numerically first, but its paired effects are only to percentage points; every Holm-adjusted survival test has , and every confidence interval excludes the prespecified five-point benefit. The ICM-by-planner interaction is also unsupported. We formalize an effective-update view in which generalized advantage estimation, episodewise standardization, PPO clipping, and mapping through policy score vectors form a many-to-one map from raw auxiliary rewards to actor updates. The experiment does not establish equal gradients or general algorithmic equivalence. It does establish a controlled case in which semantic and numerical differences among auxiliary signals fail to yield practically distinguishable downstream outcomes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.