RADIUS-DRIVE: A RISK-AWARE DIAGNOSTIC BENCHMARK IN SAFETY-CRITICAL AUTONOMOUS DRIVING
Abstract
Autonomous driving (AD) is a safety-critical domain where correct final actions may still be unsupported by reliable risk-aware reasoning. We introduce RADIUS-Drive, a diagnostic benchmark for evaluating safety-critical driving beyond outcome correctness. In particular, a model may produce a correct driving recommendation while failing to recognize the risk trigger, localize the dominant risk element, or correctly identify the factors that should govern the decision, resulting in an outcome-correct yet weakly supported decision. To support auditable evaluation under long-tail conditions, we construct a scalable scenario generation pipeline with controllable reference-based injection and reference-free generation, together with structured supervision for risk perception and decision-relevant factors. We further propose a SAR diagnostic protocol (Safety Awareness Reasoning), where S-Pass, A-Pass, and R-Pass quantify stage-wise safety, risk awareness, and decision reasoning. Complementing these stage-wise metrics, G-Pass provides a graded assessment of cross-phase decision grounding, measuring how well observable risk-aware judgments jointly support the final decision. Benchmarking 15 VLMs and four representative VLM-based AD agents reveals substantial headroom, with G-Pass ranging from to . The results further expose pronounced stage-wise disparities; for example, GPT-5.1 achieves S-Pass but only R-Pass, while the best G-Pass remains . These findings show that strong action-level safety or isolated stage performance does not necessarily translate into well-grounded safety-critical decisions, motivating diagnostic evaluation beyond final-outcome correctness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.