acceptodds
Under review as a conference paper at ICLR 2027

MechDS: Mechanistic RLVR Data Selection via Reasoning Component Dependence

Abstract

Reinforcement learning with verifiable rewards (RLVR) has become a central paradigm for improving the reasoning capabilities of large language models, yet effective data selection remains a key challenge. Existing RLVR data selection methods typically estimate the value of training data through extrinsic proxy signals, while the model's internal reasoning mechanisms remain largely unexplored. We study this missing dimension through reasoning component dependence (RCD), which measures how strongly a training sample relies on components critical to reasoning during generation. To examine information beyond extrinsic proxies, we compare RLVR training on samples with similar behavioral entropy but different reasoning component dependence. Empirically, a consistent pattern emerges across entropy-matched comparisons: low-dependence samples yield better RLVR performance than high-dependence samples, while high-dependence samples converge faster and exhibit smaller answer-reasoning gaps. These findings motivate MechDS, a dynamic data selection strategy that adaptively balances stability, learning potential, and exploration according to behavioral entropy and reasoning component dependence. Experiments on three models and three benchmarks show that MechDS consistently improves RLVR performance over strong data selection baselines, demonstrating that reasoning component dependence provides mechanistic information beyond extrinsic signals for selecting effective RLVR data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.