SIC-Drive: Sparse Instruction Conditioning for Closed-Loop End-to-End Driving
Abstract
End-to-end autonomous driving has achieved significant progress by learning a direct mapping from sensory observations to motion planning. However, future autonomous systems require not only the ability to drive autonomously, but also the ability to adapt their behaviors according to human intentions in open and dynamic environments. Natural language provides a flexible interface for specifying high-level driving objectives, such as safety-related interventions, and temporary behavioral adjustments, which are difficult to infer solely from visual observations. Despite recent advances in language-conditioned driving, incorporating language inputs does not necessarily guarantee instruction-controlled behaviors. Existing approaches often focus on integrating language representations with driving policies or generating language-conditioned trajectories, while lacking explicit mechanisms that connect human intent with planning decisions. As a result, it remains unclear whether and how language instructions actually affect closed-loop driving behaviors. In this work, we study language-conditioned controllability for autonomous driving, aiming to enable driving systems to reliably translate natural-language instructions into behavioral changes. We propose SIC-Drive, an explicit instruction grounding framework that decomposes driving instructions into lateral and longitudinal semantics and injects them into trajectory planning through path and velocity candidate scoring. Furthermore, we construct a language-augmented Bench2Drive dataset with temporally aligned driving instructions and introduce DIF-Benchmark to evaluate whether driving agents can follow instructions beyond conventional driving performance metrics. Extensive closed-loop experiments demonstrate the effectiveness of our method, highlighting the importance of explicit language grounding for controllable autonomous driving.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.