DSA: Progress Inside the Group Standardizer in Sparse-Reward Agent Learning
Abstract
Sparse terminal rewards can silence a group-relative learner even when the states inside a comparison group differ in their proximity to the goal: equal returns are mapped to equal, zero-centered coefficients. We introduce Dense Signal Augmentation (DSA), an advantage operator that combines action-before state progress with return before applying the inherited group standardizer. This placement preserves the original grouping and optimizer, recovers the baseline exactly when augmentation is disabled, and exposes progress variation without defining a second normalization scale. A conditional return identity explains the temporal alignment, while a score-function analysis identifies why current-batch scheduling, sampled grouping, and nonstationary proxies prevent an invariance or unbiased-credit guarantee. Predeclared matched training experiments exhibit earlier held-out attainment on average, although the exact paired evidence remains statistically inconclusive. Matched operator controls further test signal placement and temporal alignment, while history-aware and cross-setting experiments delimit deployment and breadth. The principal contribution is therefore a controlled and falsifiable interface between state progress and group-relative credit assignment, supported by training interventions rather than snapshot diagnostics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.