acceptodds
Under review as a conference paper at ICLR 2027

PreFTO: Reinforcement Fine-Tuning of Diffusion Planners with Local Execution Feedback

Abstract

Imitation-trained diffusion planners generate long reference trajectories, but de- ployed vehicles execute only a short segment before replanning. Reinforcement fine-tuning therefore needs feedback that connects generated plans to their long- term driving consequences. Scoring only a short execution can miss delayed hazards, whereas following an entire plan can penalize continuations that deploy- ment would replace. Full rollouts with repeated replanning are costly, and local scores must also distinguish candidates under realistic execution conditions. We study how finite local execution feedback can improve a replanning planner while retaining its pretrained trajectory generator. Under a safety-preserving continuation model with a common expected quality decrement, our analysis identifies group- centered local scores as relative advantages with an additional safety-deficit penalty. A separate sufficient condition connects score gains to deployment improvement through candidate-dependent score–value discrepancy and execution discrepancy. We introduce Prefix-Rewarded Full-Trajectory Optimization (PreFTO). It compares plans from shared policy-visited states through surrogate controller execution, uses continuous safety margins with minimum aggregation, and applies group-relative feedback to the complete original plans with DiffusionNFT. This separates the feed- back window from both the generated trajectory and the deployment replanning interval. On reactive nuPlan, PreFTO improves the pretrained Diffusion Planner on all three tested splits, including a 4.58-point gain on Test14-hard. Ablations demonstrate that proper choices of feedback length, learning targets, and loss coverage can effectively improve the closed-loop behavior of an RL post-trained planner.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.