Harnessing Learned Planners: Efficient Reinforcement Fine-Tuning with Rank-Tree Sampling
Abstract
Learned multi-modal planners can be harnessed with inference-time search: wider search can recover useful alternatives, particularly on difficult cases, but increases inference latency. In this paper, we investigate how to transfer the gains from inference-time search back into the planner itself through reinforcement fine-tuning, while keeping the cheap deployment search configuration unchanged. The challenge is constructing informative rollout groups when planning involves both discrete continuation decisions and continuous proposal geometry. Naively enumerating initial proposals with greedy continuations can miss useful later decisions, while groups with identical returns provide no learning signal under group-relative optimization. We introduce group-based rank-tree sampling, which enumerates short top-K continuation-rank schedules under a frozen reference planner, combines them with independent geometry samples, and avoids global beam pruning during training. This provides structured coverage of discrete planning decisions while preserving variation in continuous proposal geometry. We further incorporate failure-aware reward shaping and a bounded imitation-gradient contribution to stabilize learning. On NAVSIM v1, fine-tuning improves the driving planner's PDMS Score from 93.7 to 93.9 under identical evaluation settings. On an internal benchmark of dead-end parking scenarios, fine-tuning improves the success rate from 84.7% to 92.2% using a fixed beam-width-8/top-6 evaluator, exceeding the 91.4% the original model reaches only with a wider beam-width-30/top-15 search. In zero-shot evaluation on the public ParkBench dataset, the fine-tuned planner also improves the success rate from 68.6% to 72.5%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.