acceptodds
Under review as a conference paper at ICLR 2027

TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

Abstract

Retrieving scenarios of interest from large driving-video collections is essential for data curation, model development, and safety validation. Program-based systems provide explicit control over well-specified events, but commonly depend on auxiliary structured inputs and an additional language-to-program translation step. Multimodal embedding models offer a more direct interface through single-vector text–video retrieval. However, our evaluation across several pretrained embedding models reveals weak performance on driving videos, particularly for motion-centric queries. Standard supervised fine-tuning substantially narrows the domain gap but remains insufficient for capturing fine-grained driving dynamics. Ego trajectories offer a denser description of motion than captions, and their pairwise similarity can serve as graded motion supervision. Based on this insight, we introduce TraVEL (Trajectory-Guided Video Embedding Learning), a trajectory-aware adaptation framework that uses ego trajectories as privileged training supervision. TraVEL applies Group Relative Policy Optimization with complementary rewards: a graded video-to-video reward organizes the embedding space according to ego-motion similarity, while a text-to-video reward preserves language alignment. Trajectories are used only during training, and inference remains single-vector retrieval without ego poses or auxiliary perception outputs. We construct two complementary driving-video retrieval benchmarks derived from nuReasoning and the Argoverse 2 (AV2) Sensor Dataset for in-domain and zero-shot cross-domain evaluation. Across three embedding-model families, TraVEL consistently improves motion retrieval over supervised fine-tuning while maintaining comparable instance-level performance. On Qwen3-VL-Embedding-2B, it improves lateral and longitudinal motion mAP by 7.5 and 10.3 points, respectively. These improvements also transfer to zero-shot evaluation on the AV2 benchmark, demonstrating cross-dataset generalization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.