VLALIGHT: A Vision-Language-Action Model for Traffic Signal Control
Abstract
Traffic signal control (TSC) plays an important role in managing urban traffic flow and reducing congestion. In recent years, roadside cameras have become widely deployed at signalized intersections, providing continuous visual observations of traffic conditions. However, the potential of using these visual observations directly for traffic signal control remains largely unexplored. Most reinforcement learning and large language model (LLM)-based controllers still rely on complete, structured traffic states provided by simulators, which offer an idealized description of traffic conditions. Recent vision-language models (VLMs) introduce visual scene understanding, but commonly use VLMs as perception modules or high-level planners while delegating signal control to separate RL or language-model components. To bridge this gap, we present VLALight, an end-to-end vision-language-action (VLA) model for traffic signal control. Specifically, VLALight extracts traffic information from physical visual observations and reasons over dynamic scenes with multiple vehicles, lanes, and directions. It then incorporates information from neighbouring intersections to coordinate signal decisions across the road network. Besides, we introduce a cooperative reinforcement learning framework that leverages local and network-level traffic feedback to train a VLA agent that coordinates multi-intersection signal control and adapts its reasoning depth to traffic complexity. Extensive experiments on both synthetic and real-world traffic datasets demonstrate the effectiveness, scalability, and robustness of VLALight across diverse traffic scenarios.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.