acceptodds
Under review as a conference paper at ICLR 2027

Look to Understand: Small-Large Model Collaboration for Traffic Scene Understanding

Abstract

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in traffic scene understanding and visual question answering. However, large VLMs still face several challenges in complex traffic scenes, including redundant visual information, insufficient utilization of spatial information, and unclear correspondence between question semantics and relevant visual regions. To address these challenges, we propose Look to Understand, a Small-Large Model Collaborative Framework that leverages small models for visual perception and large VLMs for scene semantic understanding and visual question answering. First, we design a Small Model-Guided Perception Chain for structured visual perception. Given a traffic scene image, a lightweight vision model extracts object categories and spatial bounding boxes. The detection results are encoded as structured perception tokens, providing the large VLM with explicit object category and spatial location information. Furthermore, we propose Small Model-Guided Question-Region Contrastive Learning, which leverages the object regions provided by the lightweight vision model to model their relevance to the question semantics and aggregate their features. The question and visual region representations are then projected into a shared latent space, where bidirectional contrastive learning is employed to constrain their semantic correspondence. Experiments on RoadSceneVQA and DRIVINGVQA demonstrate the effectiveness of the proposed framework.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.