acceptodds
Under review as a conference paper at ICLR 2027

Credit Where Reasoning is Due: Learning When to Reason via Treatment-Effect-Aware Credit Assignment

Abstract

Reinforcement learning has made language models increasingly capable of reasoning, yet when to reason remains largely unaddressed. We show that standard group-relative RL cannot learn this decision when every rollout in a group reasons, as happens on-policy: the per-input effect of reasoning is absent from the policy gradient, regardless of the outcome reward. Forcing both modes into each group makes this counterfactual identifiable, but is not sufficient: the resulting signal must credit reasoning for its effect rather than for the output's quality, and it must remain reliable under noisy per-input comparisons. We address both by formulating selective reasoning as a treatment-effect learning problem and propose Treatment-Effect-Aware Credit Assignment (TECA), which combines forced 50/50 mode mixing with importance correction, a reward that scores both modes on the same scale, and a persistent empirical-Bayes estimate of whether reasoning beats direct generation on each input. That estimate gates reasoning-token credit and sets the target for the mode decision. On machine translation with Qwen3-8B, where reasoning barely moves mean quality yet widens its distribution, TECA matches or beats always-reasoning baselines while reasoning on selective inputs, and significantly outperforms the same objective without the gate. Learning when to reason therefore requires explicit treatment-effect credit, not only outcome-based RL or exposure to both modes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.