LLM-Guided Multi-Agent Reinforcement Learning for Emergency Traffic Signal Control
Abstract
Emergency vehicle (EMV) signal preemption requires coordinated multi-agent decisions under uncertain arrival information and changing traffic environments. We propose a large language models (LLM)-guided multi-agent reinforcement learning (MARL) framework integrating probabilistic perception, preference formation, and action grounding. A probabilistic long short-term memory (LSTM) estimates EMV arrival times and uncertainty, and an LLM converts this temporal context and multi-agent traffic state into phase preference distributions that modulate learned Q-values. An uncertainty-aware grounded action transformation (UGAT) reduces the impact of environment-dependent dynamics mismatch on action selection. Experiments on three real-world traffic networks show consistent improvements over MARL and traffic-control baselines. Under heavy congestion, the method reduces overall and EMV travel time by 24.96% and 44.88%, respectively, while improving traffic stability. Cross-environment experiments demonstrate robustness to changes in network structure and traffic demand.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.