SEAM: Shared-Expert Adaptation Masked by Closed-Loop Scores in Reward Fine-Tuned Diffusion Planners
Abstract
Reward fine-tuning is an appealing way to optimize safety and comfort in diffusion-based driving planners, but a higher training reward need not mean better closed-loop driving. We fine-tune the Diffusion Planner by backpropagating through the 10-step DPM-Solver++ that it runs at deployment, on two independent stacks, Autoware and nuPlan. The optimized training reward rises by a factor of 2.5, yet on the nuPlan val14 closed-loop score no fine-tuned configuration improves on the behavior-cloning checkpoint beyond seed variation. The planner is nonetheless learning. Using Shared-Expert LoRA (SE-LoRA), which adds one low-rank adapter shared by all driving domains and one adapter per domain, we show that the adaptation is structured. Domains with little training data are carried by the shared adapter, while domains with more data rely increasingly on their own. Removing the shared adapter collapses the smallest domain even though more parameters are trained, and removing the per-domain adapters costs the largest domain 15 percent. The closed-loop score misses this because the per-scenario changes are small, point both ways, and barely change how many scenarios pass or fail the 0/1 gates that the score is built from. For reward fine-tuning, this means a rising training reward is not evidence of better driving, and a flat closed-loop score is not evidence that nothing was learned.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.