CoopLight: Cooperative Vision-Language Agents for Long-Tail Traffic Control
Abstract
Traffic light signals and connected and automated vehicles (CAVs) are increasingly interdependent at urban intersections, yet they are still predominantly studied in isolation. Existing methods mainly rely on low-dimensional traffic states, limiting their ability to handle long-tail scenarios such as incidents, jaywalking pedestrians, and emergency-priority events. Such settings require semantic scene understanding, context-aware decision making, and timely cooperation between infrastructure-side and vehicle-side agents. To address these challenges, we introduce CoopLight, a vision-language framework for traffic signal and CAV control. Agent-specific visual question answering (VQA) tasks ground spatial scene understanding and decision making in multi-view observations. Multi-agent co-training further enhances agents’ ability to integrate local observations with received semantic reports. This learned cooperative reasoning enables signal and vehicle agents to use complementary evidence and adapt their decisions under partial observability. Extensive experiments on two real signalized intersections under multiple long-tail traffic scenarios show that CoopLight consistently improves traffic efficiency over rule-based, reinforcement-learning-based, and prior vision-language-model-driven baselines, demonstrating its effectiveness and proactivity for next-generation urban mobility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.