acceptodds
Under review as a conference paper at ICLR 2027

Behavioral Response Spectroscopy: Pre-RL Responses Reveal Shared Structure in RL-Induced Behavioral Change

Abstract

Reinforcement learning (RL) improves a target reward but often changes other properties of a model's responses as well. Can any of these accompanying changes be anticipated before training? We show that variation already present in the starting model contains a measurable signal of later change. Behavioral Response Spectroscopy (BRS) uses multiple answers to the same prompt to measure which response properties tend to accompany higher reward. We then compare these pre-RL directions with the behavioral changes produced by actual RL across ten models and four tasks. Our strongest test asks whether this structure transfers across objectives: we exclude the training reward from the measured directions and use separate prompts for evaluation. On two reasoning tasks, a two-dimensional space built from other reward directions captures 50.7% of post-RL displacement energy across five Qwen2.5 models and 73.2% across five other models, compared with 24.8% and 36.2% for principal components computed from the same responses without reward information. Matched controls reveal a prominent length-associated behavioral mode: multiple response properties co-vary with length, while different rewards can orient this mode toward either shorter or longer answers. This directional regularity is not confined to the original surface features: in an embedding-based behavioral representation, alignment remains positive after pre-RL length adjustment, although the relative advantage of BRS over PCA varies across representations and model cohorts. Overall, pre-RL responses already reveal a transferable component of how RL will reshape behavior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.