Is Tree RL Losing to Standard RL? Validation Mismatch in Tree-Structured RL
Abstract
Tree-structured reinforcement learning branches rollouts from shared prefixes, reusing generated tokens and turning outcome rewards into process supervision, and tree-RL methods have reported consistent gains over GRPO. Against DAPO, a more recent post-training RL algorithm that corrects the entropy collapse and instability of GRPO, these gains are no longer observed under standard validation: across three tree-RL methods, three model scales and four tasks, with the number of training trajectories per update matched, the tree-trained policy trails the DAPO-trained policy in all five comparisons, by 1.0 to 4.6 points. We identify a validation mismatch in this assessment, a mismatch between how training rollouts are generated and how checkpoints are validated: standard validation evaluates every checkpoint with greedy decoding or independent samples, the procedure that standard RL trains with but not the one tree RL trains with. Tree validation, which evaluates checkpoints with the branching procedure that tree RL trains with, corrects the assessment. Under tree validation, the tree-trained policy outperforms the DAPO-trained policy in all five comparisons, by 0.9 to 4.3 points, and it reveals training progress that standard validation misses: on Qwen2.5-Math-1.5B-Instruct, TreeRL shows no improvement over its initialization under standard validation but a gain of 4.5 points under tree validation. The same pattern holds at deployment: in every setting we test, a tree-trained policy deployed with a tree has the highest mean accuracy of the six combinations of training method and inference procedure, while generating fewer tokens than a majority vote over independent samples. Tree RL should therefore be validated and monitored with the procedure it is trained with, alongside standard validation, with the generator, answer selector and budget of each procedure specified.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.