RACS: Response Activation Consistency-guided Data Selection for Reinforcement Learning with Verifiable Rewards
Abstract
Reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning abilities of language models. Training requires many sampled responses, so selecting useful prompts is important for making effective use of the training budget. A common approach to RLVR data selection relies on prompt difficulty estimated from pass rates. Pass rate summarizes response correctness but does not capture the current policy’s behavior during generation. Unlike prior work, we study response activation consistency, the directional agreement among hidden-state vectors of multiple responses sampled from the current policy for the same prompt. In our pilot study, training on prompts with higher response activation consistency yields better test accuracy. Inspired by this, we propose Response Activation Consistency-guided Data Selection, which combines response activation consistency scoring, reference-based score estimation, and consistency-guided prompt sampling. Each epoch, it measures consistency on a small reference set containing fewer than 2% of the corpus, estimates scores for the rest, and favors high-scoring prompts for training. On five mathematical reasoning benchmarks, our method improves pooled accuracy on Qwen2.5-Math-1.5B by 3.24 percentage points over random selection and 2.92 points over the strongest evaluated baseline under the same training-update budget. Further analysis shows that high-consistency prompts tend to have more separable correct and incorrect response representations, suggesting a possible explanation for the observed gains. Our work identifies response activation consistency across sampled responses as a useful signal for RLVR data selection and provides a practical procedure for using it during training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.