Learning to Hold: Fleet-Aware and Temporally Augmented Multi-Agent Reinforcement Learning for Bus Control
Abstract
Effective bus holding control (BHC) is a real-time multi-agent decision-making problem where buses must coordinate under directional, time-varying, and partially observable fleet interactions. Existing BHC methods often rely on local observations, limited temporal context, or centralized training schemes that scale poorly with fleet size. We propose SMART-HOLD, a scalable multi-agent reinforcement learning approach for dynamic BHC. SMART-HOLD introduces a Fleet-wide State Representation that allows each bus agent to reason over fleet-level spatial structure, and a critic-only Temporally-Augmented State Representation that captures short-term system evolution for more accurate value estimation without leaking future information to the policy. A cross-attention actor-critic architecture selectively models subject, active, and passive agents to capture directional and role-dependent interactions. For scalable coordinated training, a Regionalized Training with Distilled Coordination strategy is developed to train multiple regional teacher policies in parallel and distills them into a shared actor for decentralized execution. Experiments on six real-world bus routes, including unseen routes with different service frequencies, fleet sizes, and temporal variability, show that SMART-HOLD substantially reduces bunching and improves service reliability over state-of-the-art BHC and multi-agent learning baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.