PUSH: Sparse Representation Dynamics Enable Cost-Efficient Reinforcement Learning in LLMs
Abstract
Reinforcement learning (RL) for large language models relies on costly on-policy sampling. Exploiting the representational trajectory traversed during RL is a promising direction for improving training efficiency and model performance. We use sparse autoencoders (SAEs) to decompose an LLM's residual-stream activations into high-dimensional feature coordinates to interpret the model state during the RL process. The token-aligned change in SAE feature activations from the pre-RL model to the current model forms the SAE signal. We find that the direction along which the model improves during RL is low-dimensional in SAE coordinates, with five features at the final transformer layer accounting for 96.83% of the aggregate squared norm in code generation. These directions persist throughout training so that the signal predicts where RL is heading despite non-monotonic training dynamics. We introduce Progress Utilization via Sparse Hidden Representations (PUSH). On a fixed set of guidance texts, PUSH resolves this signal into its dynamic features and extends their per-token targets along the directions these features have established over training. Each PUSH application fits these targets in a minute-scale activation-matching update outside the RL reward channel and standard RL then resumes under its original objective. Experiments in code generation and mathematical reasoning show immediate performance gains from PUSH without additional RL rollouts, and these gains persist during continued RL. Applying PUSH repeatedly reduces the cumulative sampling cost of RL by 30% on average at matched accuracy. Our results show that PUSH uses interpretability to explain prior learning and direct subsequent training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.