LTDriveQA: Benchmarking Spatiotemporal Intelligence in Multi-View Long-Tail Driving Scenarios
Abstract
Driving visual question answering (VQA) provides a unified framework for assessing whether vision-language models (VLMs) and Vision-Language-Action (VLA) models can understand dynamic road scenes and reason about driving decisions. However, existing benchmarks often rely on generic templates, weakly probe temporal and cross-view reasoning, and underutilize long-tail driving events. We introduce LTDriveQA, an event-centered benchmark of scene-specific questions designed to probe temporal and cross-view reasoning. It spans Perception, Prediction, Planning, and Reasoning, supports both multiple-choice and open-ended evaluation. A scalable recursive benchmark construction pipeline refines questions through diagnostic feedback, evidence grounding, answer annotation, and restricted-input shortcut checks. Across 25 models, GPT-5.6-Sol leads with an overall score of 60.0; Qwen3-VL-8B and Qwen-Drive-1.0 score 40.5 and 45.6, respectively. Evidence ablations show aggregate visual gains across all open-source models, with temporal gains strongest in Perception and surround-view gains consistent across three models on multi-view questions. By grounding evaluation in challenging real-world driving, LTDriveQA probes the scene understanding and decision reasoning that driving VLMs and VLAs rely on, and provides a testbed for spatiotemporal intelligence in dynamic multi-view environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.