acceptodds
Under review as a conference paper at ICLR 2027

ISO: An RLVR-Native Optimization Stack

Abstract

Reinforcement learning with verifiable rewards (RLVR) advances the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. We study this layer and functionally validate spectral inheritance: RLVR can reuse the base model's weight spectra while acquiring new behavior by adapting both the associated input and output singular frames. We operationalize this finding as Isospectral Optimization (ISO), an RLVR-native optimization stack with complementary offline and online instantiations. Offline, ISO-Merger composes shared-base specialists in fixed-spectrum frame coordinates without post-merge data, rollouts, or distillation, and achieves the highest aggregate scores among the data-free merging methods. Online, ISO-Optimizer applies AdamW or Muon directly to the frame variables under fixed base spectra, improving aggregate accuracy in the reported math and coding experiments. On Qwen3-4B-Base, ISO-AdamW reaches AdamW's 400-update accuracy in 280 updates with 1.5 X fewer updates and 27% less wall-clock time, and exceeds AdamW by 2.2 points at update 400. These results support a reusable RLVR design principle: inherit the spectrum, optimize the frames.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.