Seeing with Tools, Reasoning with Language: Recasting Vision as Structured Scene Text for Driving VQA
Abstract
Driving visual question answering (VQA) requires extracting relevant visual information from images or videos and reasoning over it. Current vision-language models (VLMs) couple these tasks, obscuring whether their errors arise from visual information extraction or downstream reasoning. We investigate this distinction with DriveTLM, a training-free framework that delegates visual information extraction to visual processing tools, combining off-the-shelf vision models with deterministic algorithms. Their outputs are organized into structured scene text, an explicit language interface that lets the answering model reason without visual tokens. The framework supports images and videos, Rule-based or Agentic tool selection, and both VLM and LLM answer backbones without parameter updates. Across image benchmark DriveBench and video benchmark LingoQA, DriveTLM improves over direct visual input for multiple model families and sizes, with particularly strong gains for smaller backbones and competitive end-to-end latency. Controlled studies examine these gains: most persist after removing rule-derived risk and answer-adjacent statements, expressing the same information as text is more effective than visual overlays, and selecting relevant information matters more than including every available detail. The model also responds selectively when individual facts are changed. These findings show the value of structured scene text as an interface for delivering tool-derived evidence to frozen answering models. The code will be released upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.