acceptodds
Under review as a conference paper at ICLR 2027

When Does Branching Pay? An Exact Variance Analysis of Tree-Structured Policy Gradients

Abstract

Group-relative policy gradient methods sample a fixed number of independent completions per prompt, so the rollouts share nothing but the prompt. A growing literature makes use of tree-structured policy gradient estimators instead, branching at chosen points guided by heuristics. We formalize such designs as a pair : a kernel that decides where and how much to branch, and an anchor map that decides which ancestor supplies each edge's baseline group. For any such pair, we give an exactly unbiased on-policy gradient estimator and, to our knowledge, the first exact variance decomposition for tree policy gradients with adaptive branching and group baselines. From it, we derive design principles for where to branch and how to form baseline groups, and identify regimes in which branching does not help. We test these predictions on reasoning and multi-turn agentic tool-use tasks with verifiable rewards.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.