acceptodds
Under review as a conference paper at ICLR 2027

STRIDE: Evaluating Spatiotemporal Reasoning in Driving Edge Cases

Abstract

Evaluating vision-language models (VLMs) for autonomous driving requires assessing their ability to connect spatial observations of past events with future consequences. Existing benchmarks either assess driving edge cases through basic perceptual questions or evaluate temporal understanding under relatively simple scenarios, leaving the spatiotemporal reasoning required in complex driving scenarios insufficiently tested. To address this gap, we introduce STRIDE (SpatioTemporal Reasoning In Driving Edge Cases), an object-centric benchmark for evaluating spatiotemporal reasoning in long-tail and safety-critical scenarios. STRIDE defines six task families: spatial perception, spatial understanding, temporal memory, temporal extrapolation, trajectory prediction, and scene-context awareness. It pairs multiple-choice questions covering a broad range of spatiotemporal judgments with related free-form visual question answering (VQA) that examines the same capabilities through detailed scene explanations and justifications of driving decisions. We curate STRIDE by prioritizing scenes with unusual object behaviors and complex traffic interactions, constructing reference answers from structured dataset annotations through automated procedures, and applying limited language model assistance alongside human verification to support annotation quality. Evaluation of 12 proprietary, open-weight, and driving-specialized models reveals substantial limitations: even GPT-6-Astra achieves the highest average multiple-choice accuracy of 43.4%, compared with a random-choice baseline of approximately 20%. Our analysis further reveals that strong scene-description scores can mask weaknesses in safety-critical judgments, and models struggle more with reconstructing object motion than recalling past positions. These insights highlight STRIDE's diagnostic value for uncovering spatiotemporal driving limitations that aggregate VQA performance can conceal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.