Sequential Group-Relative Test-Time Adaptation for LLM-planner on a Closed-loop Single Driving Timeline
Abstract
LLM-based motion planners still lose substantial closed-loop performance in unseen driving domains, motivating online test-time adaptation from reward feedback. Two structural obstacles stand in the way: per-step gradient updates on a multi-billion-parameter policy are expensive, and critic-free group-relative methods such as GRPO normalize each rollout against other rollouts from the same state, which a deployed planner never observes on its single, irreversible driving timeline. We propose TeTO (Test-time Trajectory-level Optimization). Its core, Sequential Grouping, reformulates group-relative optimization along time, normalizing each rollout against the planner's immediately preceding rollouts in short, non-overlapping windows. The deviation of this baseline from the unobservable same-state baseline is bounded by the local reward drift, which is small under local smoothness, plus a sampling term; empirically, a three-step window matches the reward variance of same-observation groups, and its advantages agree in sign with a same-state reference in 92.3% of logged windows without abrupt time-to-collision changes. Reward Truncation further restricts updates to windows whose weighted per-component minimum reward falls below a threshold, and only LoRA adapters are updated, with the frozen backbone serving as the KL reference. On nuPlan, TeTO improves the supervised planner by up to 16.7% in an out-of-domain city while updating on as few as 7.8% of steps, with consistent gains across two LLM backbones, three unseen cities, and a planner trained on mixed nuPlan–Waymo data. To our knowledge, TeTO is the first critic-free group-relative policy optimization method for a single executed driving timeline whose baseline has a bounded deviation from the same-state baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.