acceptodds
Under review as a conference paper at ICLR 2027

Recon4Drive: Selective and Explicit 4D Priors for VLA Driving

Abstract

Vision-language-action (VLA) models reason effectively about driving scenes, but their image-centric representations do not explicitly capture the metric geometry and temporal dynamics required for planning. Pretrained 4D reconstruction models encode such information, yet their dense representations are designed to describe the entire scene. For planning, however, only a small portion of this information is relevant, and the current decision mainly depends on a few ego-centric relations, such as the distance to a lane boundary or the gap to a nearby vehicle, which remain implicit in dense features. To this end, we present Recon4Drive, a VLA framework that converts dense 4D reconstruction features into compact and explicit inputs for planning. The Semantic-Guided 4D Retriever (SG4R) uses image semantics to retrieve planning-relevant geometry and motion cues from the 4D representation, keeping the number of 4D tokens small regardless of its density. Planning-Critical Structured Scene Encoding (PSE) further assigns supervised queries to key relations with lanes and surrounding agents, including their future motion, making these relations directly accessible to reasoning and planning. On the closed-loop Bench2Drive benchmark, Recon4Drive achieves a Success Rate of 62.27% and a Driving Score of 85.45 with camera-only inference, the best results among methods trained on the official Think2Drive-collected data. The code will be made publicly available upon publication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.