acceptodds
Under review as a conference paper at ICLR 2027

CoopLight: Cooperative Vision-Language Agents for Long-Tail Traffic Control

Abstract

Traffic light signals and connected and automated vehicles (CAVs) are increasingly interdependent at urban intersections, yet they are still predominantly studied in isolation. Existing methods mainly rely on low-dimensional traffic states, limiting their ability to handle long-tail scenarios such as incidents, jaywalking pedestrians, and emergency-priority events. Such settings require semantic scene understanding, context-aware decision making, and timely cooperation between infrastructure-side and vehicle-side agents. To address these challenges, we introduce CoopLight, a vision-language framework for traffic signal and CAV control. Agent-specific visual question answering (VQA) tasks ground spatial scene understanding and decision making in multi-view observations. Multi-agent co-training further enhances agents’ ability to integrate local observations with received semantic reports. This learned cooperative reasoning enables signal and vehicle agents to use complementary evidence and adapt their decisions under partial observability. Extensive experiments on two real signalized intersections under multiple long-tail traffic scenarios show that CoopLight consistently improves traffic efficiency over rule-based, reinforcement-learning-based, and prior vision-language-model-driven baselines, demonstrating its effectiveness and proactivity for next-generation urban mobility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.