acceptodds
Under review as a conference paper at ICLR 2027

DriveScout: Sharpening MLLM Perception in Autonomous Driving with Adaptive and Robust Visual Focusing

Abstract

Multimodal Large Language Models (MLLMs) are increasingly used in autonomous driving for perception, reasoning, and interpretable scene understanding. Yet safety-critical small-scale targets, such as pedestrians, traffic signs, and distant vehicles, remain difficult to recognize, particularly when motion blur weakens their already limited visual evidence. Think-with-Images MLLMs can acquire additional evidence by cropping and enlarging local regions, but blur may destabilize the selected region, while indiscriminate image operations add inference cost. To study this problem, we introduce SCOUT, a paired driving VQA benchmark built from driving scenes in nuScenes, where each clear observation is matched with motion-degraded variants for small-target understanding. We further propose DriveScout, a reinforcement learning framework that formulates visual focusing as selective evidence acquisition. DriveScout combines Target-Grounded Spatial Credit with Necessity-Aware Focus Optimization, while a Progressive Blur Curriculum supports learning as the visual evidence becomes increasingly degraded. Experiments on SCOUT and two additional driving benchmarks show that DriveScout improves robustness to blur while reducing redundant visual operations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.