acceptodds
Under review as a conference paper at ICLR 2027

Beyond Visual Realism: Benchmarking Generated Driving Scenarios as Policy Tests

Abstract

End-to-end autonomous vehicles need to be tested on rare, safety-critical situations that real-world driving data collects too slowly and expensively. Generative models offer a scalable fix—synthesizing unbounded, controllable traffic scenarios from recorded histories—but the community still lacks a standard for whether a generated scenario actually works as a driving test. Existing benchmarks assess visual fidelity, temporal consistency, or traffic plausibility, but rarely ask whether it exposes meaningful policy failures, whereas visual corruption naturally lowers policy performance. We introduce **DriveAudit**, a hierarchical benchmark for evaluating generated driving scenarios as policy tests. The multi-source dataset includes 550 scenarios, combining native multi-camera histories with monocular-derived multi-view inputs constructed through generative view completion. Evaluation proceeds through four coupled levels—observation quality, scenario fidelity, policy response, and quality-aware testing effectiveness. Observation quality assesses visual realism and temporal coherence. Scenario fidelity measures image–motion agreement and traffic plausibility through Trajectory–Visual Consistency (TVC) and Traffic Realism Score (TRS). Policy response use Responsibility-aware PDMS (RPDMS), which gates collisions by event-aligned ego attribution, and the headline Quality-aware Testing Effectiveness Score (QTES) multiplies paired policy degradation by reference-relative quality retention, so a visually corrupted scenario cannot be rewarded merely for breaking the policy. Benchmarking seven scenario-generating baselines with six driving policies reveals a mismatch between visual fidelity and testing effectiveness: models that produce more realistic observations do not consistently yield more effective policy tests. DriveAudit thus enable assessing the generated scenarios through joint analysis of visual realism, traffic fidelity, and policy behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.