TAPER: Teacher-guided, Adaptive-length Policy optimization for Efficient Reasoning
Abstract
Explicit multi-step reasoning improves the capability and reliability of large language models, but at considerable computational cost: traces are often dominated by redundant verification and unproductive backtracking. Because useful computation is entangled with this redundancy, indiscriminate shortening sacrifices both; what can be safely removed depends on how reliably the model solves the problem. We present **TAPER (Teacher-guided, Adaptive-length Policy Optimization for Efficient Reasoning)**, an on-policy framework that treats how much a model reasons and how it does so as a single learning problem. At the trajectory level, an accuracy-adaptive relative-length objective gauges the model's current competence on each problem from its own online samples, adjusting compression strength accordingly. At the token level, region-aware asymmetric teacher supervision lets a stronger teacher guide the student's own generation, with distinct objectives for the reasoning and answer regions; the two objectives are optimized jointly on the same trajectories in each update. Experiments on Qwen3.5-4B and Qwen3.5-9B across five benchmarks spanning mathematical, knowledge-intensive, and logical reasoning, against six baselines, show that TAPER improves accuracy over GRPO by 1.2–7.9 percentage points while using 18.4%–66.4% fewer tokens. Behavioral analysis further shows that TAPER reshapes the model's reasoning behavior in a difficulty-dependent manner: it reallocates the reasoning budget toward productive advancement and away from unnecessary redundancy, with the proportion of the two adjusted according to the model's competence on each problem.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.