acceptodds
Under review as a conference paper at ICLR 2027

ARDrive: Adaptive Retrieval-Augmented Generation for Autonomous Driving

Abstract

Retrieval-augmented generation (RAG) has recently been introduced to enhance the reasoning ability of vision-language models (VLMs) for autonomous driving (AD) by incorporating external driving knowledge. However, existing RAG-based AD systems typically perform retrieval indiscriminately, assuming that external knowledge is always beneficial. This assumption does not always hold. RAG can be redundant in simple scenarios and may even introduce misleading information when retrieval is imperfect, leading to extra latency and hallucinated reasoning. To address this limitation, we propose ARDrive, a RAG-aware driving framework that formulates retrieval usage as a learnable, context-dependent decision rather than a fixed inference procedure. ARDrive introduces a learnable [RAG] token as a dynamic retrieval gate and trains it through a two-stage RAG-aware pipeline. First, supervised fine-tuning (SFT) establishes initial retrieval awareness using performance-gain-driven pseudo labels generated by comparing RAG and No-RAG reasoning outcomes with an LLM-based semantic scorer. Second, RAG-aware reinforcement learning (RL) further refines the retrieval policy by directly optimizing the trade-off between reasoning accuracy and retrieval cost. Experiments on public AD benchmarks show that ARDrive reduces reasoning latency by 14% and RAG activation by 34%, while improving driving reasoning performance over conventional always-RAG baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.