VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control
Abstract
Traffic signal control (TSC) is essential for mitigating urban congestion. Recent advances in vision-language models (VLMs) enable richer interpretation of intersection scenes, opening new opportunities for visual-context-aware TSC. However, existing VLM-based controllers follow a cascaded perceive–reason–control pipeline whose vision-to-text conversion loses fine-grained spatial detail and whose sequential multi-agent inference adds latency that is problematic for real-time control. We propose VLALight, a lightweight vision-language-action (VLA) framework that maps a stitched multi-directional observation of an intersection, together with a textual description of its phase configuration, directly to discrete signal-phase actions in a single forward pass, avoiding intermediate image-to-text descriptions and handcrafted traffic-state representations. A simulation-grounded pipeline anchored to real intersection topologies provides expert demonstrations, and a compact 0.5B-parameter multimodal backbone, adapted with LoRA and a small set of learnable phase queries, answers each decision in 0.22 s on a local GPU. In closed-loop simulation on six replicas of real intersections, VLALight attains the strongest emergency-vehicle service among the compared controllers, reducing pooled emergency waiting time by 21.1% over the cascaded cloud-scale VLMLight, and generalizes to unseen intersection topologies and traffic-flow.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.