CEDAR: Concise, Evidence-grounded Driving Action Recommendation via Cascade Reasoning and Policy Refinement
Abstract
Large vision-language models have advanced autonomous driving toward interpretable scene understanding, yet existing models tend to produce verbose responses without explicitly surfacing the visual evidence behind their decisions, limiting their practicality for real-time driving assistance. We formulate the task of concise, evidence-grounded driving recommendation generation, where actionable decisions are explicitly backed by decision-relevant visual evidence and contextual reasoning. To this end, we construct CODA-GRD (Grounding, Reasoning & Decision), a unified Evidence-Reason-Suggestion benchmark providing structured supervision for visual grounding, reasoning, and action recommendation. Built upon it, we propose CEDAR, trained via a progressive pipeline: Entity Presence Grounding Pretraining aligns visual entities with structured evidence; Chain-of-Insight Cascade Reasoning (CCR) propagates context from Reason to Evidence to Suggestion under a reason-aware contrastive objective; and Driving Policy Refinement applies branch-wise Group Relative Policy Optimization with task-specific rewards. Together, these stages enable CEDAR to generate concise recommendations explicitly grounded in visual evidence and reasoning. Experiments on CODA-GRD and three external benchmarks, OmniDrive, LingoQA, and DriveCoT, show consistent improvements over strong baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.