acceptodds
Under review as a conference paper at ICLR 2027

When Format Rewards Fade: Layer-Specific Representation Shifts in RLVR

Abstract

Format constraints in reinforcement learning with verifiable rewards (RLVR) are typically treated as output-level scaffolding, but whether they leave persistent traces in a model’s internal representations remains unclear. We study this question through a controlled two-phase training protocol for mathematical reasoning. All models first undergo a shared format-compliance phase and then continue training under three conditions: format reward withdrawal, sustained fixed-format reward, or sustained randomized-format reward. Using paired-format evaluation traces and layer-wise representation analysis relative to the shared Phase-1 checkpoint, we examine how these reward choices reshape format- and problem-related information. We find that withdrawing format reward produces a pronounced increase in problem-related variance at a late-middle layer, while the corresponding changes are not reliably observed at early or final layers. Sustaining fixed-format reward instead leaves representations closer to the shared baseline, revealing a strongly layer-specific response to reward composition. Behavioral evaluation further shows that the randomized-format condition attains the highest observed mean GSM8K accuracy and the lowest cross-seed variability among the three conditions. Together, these results suggest that format-induced specialization is not necessarily permanent: changing the reward signal can selectively reshape internal representations, while format diversity may offer a promising direction for more stable RLVR training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.