acceptodds
Under review as a conference paper at ICLR 2027

Learn to Reason Efficiently on Real Paths via Tree-Structured Policy Optimization

Abstract

Large reasoning models often waste substantial test-time compute, and recent length-aware reinforcement learning addresses this issue by rewarding short and correct trajectories. We reveal a previously overlooked training–inference mismatch in the prevailing trajectory-level paradigm for length-aware RL: positive learning signals are concentrated on trajectories whose early reasoning is already well directed, whereas inference requires the model to continue from its own potentially erroneous or exploratory prefixes. Through prefix-level continuation analysis, we empirically verify this mismatch and show that conditioning reinforcement learning on self-generated prefixes substantially improves both reasoning accuracy and efficiency. To address this issue, we propose **T**ree-structured **L**ength-aware **P**olicy **O**ptimization (**TLPO**), which performs prefix-conditioned credit assignment by comparing continuations branching from the same self-generated prefix and optimizes their efficiency with segment-level length-aware rewards. Across three representative reasoning models and five benchmarks, TLPO reduces average reasoning length by 56.0% to 80.3% while improving average accuracy by 0.5 to 4.1 percentage points, with out-of-domain evaluations showing consistent gains on scientific reasoning and coding tasks. These results establish prefix-conditioned optimization over self-generated reasoning paths as an effective principle for training efficient reasoning models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.