ETM: Event Trigger Module based VLM for Autonomous Driving
Abstract
Large Vision Language Models (VLMs) have recently demonstrated remarkable semantic reasoning capabilities for autonomous driving and are increasingly incorporated into fast-slow planning frameworks. However, existing approaches invoke the VLM at every planning step even though high-level semantic changes are typically sparse over time, leading to substantial redundant computation. This raises a fundamental question: when should the expensive VLM be invoked? To address this problem, we propose the Event Trigger Module (ETM), a lightweight event-driven policy that selectively activates the VLM only when refreshed semantic reasoning is expected to improve downstream planning. We formulate VLM invocation as an information-gain-driven decision that explicitly balances semantic utility against computational cost. Guided by this formulation, ETM integrates short-term visual dynamics, long-term temporal context, and compact motion-aware scalar cues to estimate the necessity of VLM invocation. To train the trigger policy effectively, we further introduce a hybrid supervision strategy that combines corner-case trigger labels with information gain based pseudo-labels, enabling ETM to capture both common driving situations and rare but safety-critical events. Extensive experiments on the nuScenes benchmark demonstrate that the proposed ETM method achieves comparable planning accuracy to always-triggered VLM systems while reducing the VLM invocation rate from 100% to 32%. Meanwhile, ETM decreases inference latency by nearly 50% and lowers computational cost by approximately 64%, without sacrificing planning performance. These results show that the proposed ETM method offers an effective paradigm for VLM-assisted autonomous driving.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.