LongBusBench: Benchmarking LLM Agents for Long-Horizon Bus Control
Abstract
Large language model (LLM) agents are increasingly capable of long-horizon control, yet their ability to operate continuously in real-world dynamic systems remains largely unexplored. Public transit provides a challenging setting in which control actions have delayed and interacting consequences: for example, bus holding changes vehicle headways and passenger queues, thereby altering the conditions encountered by subsequent vehicles and the future validity of past experience. We introduce LongBusBench, a data-grounded benchmark for evaluating LLM agents in continual and asynchronous bus control under evolving operating conditions. We further propose MEMBus, a memory-augmented LLM agent that structures experience into episodic and procedural memories and selectively retrieves, consolidates, revises, and forgets information towards efficient transit control. Experiments across various operating protocols compare MEMBus with memory-free LLM agents, alternative memory designs, and classical control baselines. Results show that operational memory can improve public transit efficiency, while diagnostic analyses reveal that its effectiveness depends critically on the relevance and validity of accumulated experience. These findings establish long-horizon public-transit control as a practical and promising testbed for studying continual LLM agents in persistent real-world operational systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.