Modulating Language Model Reinforcement Learning with Internal Dynamics
Abstract
Reinforcement learning (RL) has become a key approach for improving large language models (LLMs), using feedback from human preferences, verifiable rewards, model signals, or environment interactions. Existing methods mainly determine whether a behavior should be reinforced or suppressed, but largely ignore how strongly the model should learn from it. We find that token representations differ substantially in how they evolve across layers: some stabilize early, while others continue to change until the final layers. These internal dynamics do not indicate whether a decision is good or bad; instead, they provide a natural signal for modulating the strength of policy updates. Based on this observation, we introduce Internal Dynamics Modulation (IDM), a lightweight operator that uses cross-layer representation dynamics to modulate the magnitude of existing policy advantages while preserving their sign. IDM is lightweight and sign-preserving, and can be inserted after an existing advantage estimator without changing the underlying reward source or policy optimization framework. Across different RL scenarios, IDM consistently improves the corresponding baselines, with larger gains on harder and longer-horizon settings. At larger model scales, internal dynamics also identify critical positions more reliably than entropy or log-probability. These results suggest that internal dynamics provide a general modulation signal complementary to the source of reinforcement feedback, offering a simple way to modulate the strength of policy updates without changing their sign.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.