Benchmarking Evidence Sufficiency in Vision-Language Models for Autonomous Driving
Abstract
Vision–language models (VLMs) for autonomous driving are typically evaluated by answer correctness, even when the supplied observations may not contain sufficient evidence to justify an answer. We introduce DriveAnswer, a benchmark of 20,151 questions across seven reasoning families, in which each question is evaluated under matched strong- and weak-evidence configurations. Each pair preserves the underlying scene fact and semantic answer while changing whether the observations support answering, making abstention the appropriate response under weak evidence. Across five zero-shot VLMs, strong-condition accuracy substantially overstates evidence-sensitive performance: the best baseline reaches 64.69% strong-evidence accuracy but only 41.96% matched-pair success. A controlled adaptation study separates answer, sufficiency, grounding, and matched-pair supervision. Answer-only fine-tuning nearly eliminates abstention, whereas strong-weak sufficiency supervision restores selective behavior; grounding supervision produces a separate localization gain. On the held-out test set, DriveAnswer-Gemma reaches 81.58% matched-pair success, while explicit grounding and sufficiency supervision raises joint localization-presence ([email protected]) from 53.43% to 77.16% and joint answer-grounding accuracy ([email protected]) to 61.98%. These results show that semantic correctness, evidence sufficiency, and visual grounding are distinct capabilities that should be evaluated jointly.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.