acceptodds
Under review as a conference paper at ICLR 2027

PASP: Positive-Anchored Spectral Projection for RLVR

Abstract

Muon has improved large language model pretraining, and Pion extends its spectral approach to reinforcement learning with verifiable rewards (RLVR). We identify a spectral asymmetry: positive-advantage gradients concentrate in a low-dimensional subspace and have lower effective rank than negative-advantage gradients. Muon and Pion transform these gradients after combining them into momentum, without separately controlling the two channels. This motivates PASP (Positive-Anchored Spectral Projection), which uses the current positive gradient to define an anchor subspace at each step. PASP jointly transforms the full positive gradient and the part of the negative gradient within this subspace. It separately transforms a low-rank approximation of the remaining negative gradient and adds the two outputs to form the update. We evaluate PASP on Qwen3 models against AdamW, Muon (hidden-only), and Pion. PASP achieves the highest mean accuracy across four reasoning benchmarks on both Qwen3 models. On Qwen3-8B, it reaches 0.489 ± 0.006, compared with 0.470 ± 0.015 for Pion, where ± denotes the sample standard deviation over three training seeds. On Qwen3-1.7B, the 8B recipes are transferred without retuning. All three PASP runs complete training, averaging 0.463, compared with 0.398 for Muon and 0.184 for Pion. Under DAPO, PASP reaches 0.485, compared with 0.466 for Muon and 0.468 for Pion. It also remains competitive on GPQA-Diamond and AMC23-Chain4, although the gains vary across settings. These results support using advantage sign to guide spectral optimization in the tested RLVR settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.