AttendingScorer: An End-to-End Driving Evaluator that Read Out Trajectory Evaluation Score from latent of Vision-Language-Model
Abstract
Driving vision-language models are now capable of acting as effective evalua- tors between competing trajectories, and providing accurate explanations for their choice. However, they still struggle to provide accurate quantitative scores to these on a non-comparative scale. Existing learned evaluators address this gap by either relying on predefined heuristic rules, or through techniques that involve finetuning the existing comparison-based evaluator to also provide a numerical score. The former approach struggles with unforeseen, long tail events, while the latter involves repeated modification of model weights which can have uninten- tional knock on effects on existing evaluator capability. We propose TrajSteering, a learned approach that aims to read out the trajectory score from the evaluator, AttendingScorer, without touching its weights. It trains only a small set of query embeddings, including a designated score query that aims to get a score through a sum of the context values, weighted by attention, as dictated by the other query embeddings. Our key architectural observation is that attention weights are thus analagous to classical filter taps, with the key distinction being that with filter taps the weights can take on any values, attention weights are restricted to what the frozen model can actually realize. Thus they can be approached with similar prin- ciples. TrajSteering combines two objectives. An automatic pseudo-score loss in- tends to train a scalar head to extract a numeric score from the score-query output, to match targets generated offline by a rule-gated teacher. A deflection-criterion, conditioned on the same teacher, identifies which query embeddings best isolate score-relevant signal from irrelevant noise, with the aim of selecting a readout ca- pable of generalizing among the ones that fit. The backbone remains unchanged; the teacher is used only for one-time offline supervision. Using extensive experi- ments, AttendingScorer perform at a high level of evaluation quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.