acceptodds
Under review as a conference paper at ICLR 2027

CAST: Credit Assignment on Selective Reasoning Trees for Long-Chain Reasoning

Abstract

Reinforcement learning with verifiable rewards has become the standard recipe for improving long-chain reasoning. However, current methods struggle to put compute and training signal where they matter most in a trajectory. On hard problems, an early mistake often derails the trajectory, and later reflection rarely fixes it. On easier problems, once the model has enough information to determine the correct answer, it keeps generating redundant verification and self-checks. Consequently, training signals get diluted across uninformative steps, and rollout compute is wasted on reasoning that no longer changes the outcome. To resolve both bottlenecks, we propose Credit Assignment on Selective Reasoning Trees (CAST), a tree-structured RL framework for long-chain reasoning. During rollout, CAST builds selective reasoning trees, expanding only at unresolved branch points where sampled continuations yield mixed outcomes (both correct and incorrect). It then pools all leaf rewards and the tree topology to estimate intermediate step values in closed form, deriving fine-grained edge advantages. Across five diverse backbones, ranging from 1.7B base models to a 35B mixture-of-experts, CAST consistently improves accuracy while generating shorter responses. On Qwen3-4B-Base, CAST outperforms DAPO by 4.03 points on average, cuts response length by 25.8%, and matches AIME24 accuracy with a 1.8× training speedup. By focusing rollouts on decisive steps and sharing credit across the tree, CAST shows that targeted exploration and structured credit assignment make long-chain reasoning both stronger and more efficient. Upon acceptance, we will release our code and model checkpoints.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.