Beyond Endpoint Flatness: Policy Scale Dynamics Shape Robustness in SAM-PPO
Abstract
Sharpness-Aware Minimization (SAM) improves PPO robustness to action noise, and prior work links this gain to flatter endpoint reward landscapes. We study how this robustness forms and what endpoint geometry can identify in diagonal-Gaussian PPO with learnable training scale, evaluating deterministic mean policies under external action noise. Across four MuJoCo tasks, replaying complete scale histories shows that action robustness largely follows the training-scale history, while changing the mean updater has smaller, task-dependent effects. Local analysis and routing interventions further support feedback from mean perturbations into the scale update, and explicit scale controls show that this pathway can be regulated without the full SAM update. With training scale fixed, Mean-SAM produces flatter relative reward landscapes without consistent action-noise gains. Equalizing the absolute parameter perturbation dose substantially weakens this flatness advantage, and further matching the induced action displacements reveals no consistent return-retention advantage across tasks. Together, these results identify training scale as a functional pathway behind SAM-PPO robustness and show that endpoint flatness alone cannot determine how that robustness was formed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.