acceptodds
Under review as a conference paper at ICLR 2027

Seeing with Tools, Reasoning with Language: Recasting Vision as Structured Scene Text for Driving VQA

Abstract

Driving visual question answering (VQA) requires extracting relevant visual information from images or videos and reasoning over it. Current vision-language models (VLMs) couple these tasks, obscuring whether their errors arise from visual information extraction or downstream reasoning. We investigate this distinction with DriveTLM, a training-free framework that delegates visual information extraction to visual processing tools, combining off-the-shelf vision models with deterministic algorithms. Their outputs are organized into structured scene text, an explicit language interface that lets the answering model reason without visual tokens. The framework supports images and videos, Rule-based or Agentic tool selection, and both VLM and LLM answer backbones without parameter updates. Across image benchmark DriveBench and video benchmark LingoQA, DriveTLM improves over direct visual input for multiple model families and sizes, with particularly strong gains for smaller backbones and competitive end-to-end latency. Controlled studies examine these gains: most persist after removing rule-derived risk and answer-adjacent statements, expressing the same information as text is more effective than visual overlays, and selecting relevant information matters more than including every available detail. The model also responds selectively when individual facts are changed. These findings show the value of structured scene text as an interface for delivering tool-derived evidence to frozen answering models. The code will be released upon publication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.