OPTS-TTPO: Enhancing Finite-Sample Policy Gradient Learning with Tree Search
Abstract
The policy-gradient theorem expresses the exact gradient as an expectation under the current policy, but practical methods estimate it from finitely many on-policy trajectories. Rare high-return trajectories may be missing from the policy-gradient estimate. We study whether tree search can improve their coverage under a finite budget while controlling the induced gradient bias. We introduce On-Policy Parallel Tree Search (OPTS) and Tree Trajectory Policy Optimization (TTPO) through on-policy tree trajectories, which sample new suffixes from the current policy at visited states. This requires no action-distribution correction, but branching changes state visitation. Our Branch Aggregation Lemma underpins TTPO: branch-weighted tree statistics recover chain expectations when branch choices and weights are fixed before outgoing transitions are sampled. OPTS selects expansion states using estimated performance differences. Under deterministic dynamics, exact values, and max-backup advantages, the induced search policy's expected return improves monotonically with the budget. Adaptive expansion violates the lemma's condition; we bound its gradient bias and show that max backup adds prefix credit to actions leading to better discovered suffixes. Against a finite chain reference, TTPG's measured bias stays near its no-branching level while NaivePG's bias grows from to . Exact-value search is monotone and learned-critic summaries improve in aggregate. At matched budgets, reward- and value-guided OPTS improve correct-answer coverage and majority-vote accuracy over i.i.d. baselines. The coverage–bias diagnostic shows a larger gain from OPTS + TTPG at a smaller bias shift than Fixed-branch + NaivePG. Under matched interaction or rollout budgets, OPTS-TTPO improves MuJoCo tail returns over PPO by up to , raises Atari-57 human-normalized IQM from to , and improves micro-averaged pass@32 across all four Qwen3 models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.