TSCraft: An Experience-Guided Agent for Evolving Traffic Signal Control Policies
Abstract
Traffic signal policies must adapt to demand while remaining fast and auditable. Rule-based policies need manual revision, deep RL encodes control in learned parameters, and online language-model controllers generate and validate actions inside the control loop. We present TSCraft, an experience-guided agent that evolves compact code policies offline and freezes one for deployment. It maintains multiple policy lineages and revision experience, evaluates anonymous candidates on matched simulator scenes, and maps lane-group demand to legal phases through a topology-normalized interface. The frozen policy runs locally at each intersection without language-model calls, parameter updates, or inter-intersection communication. In one 20-round campaign, TSCraft evaluated 40 candidates and promoted eight distinct non-seed policies, producing an explicit stepwise trajectory. Across matched flows and horizons, each controller is evaluated through its released deployment interface. Under this end-to-end protocol, TSCraft achieves 15.30 s mean waiting on a 23-scenario test constructed after policy selection, versus 25.08 s for MaxPressure, with lower waiting in 22 scenarios. When the same source is deployed at every intersection, it reduces average waiting time relative to MaxPressure by 25.49%, 20.96%, and 22.50% on the evaluated Hangzhou, Jinan, and New York flows. Across 23 sealed episodes, a replay record reconstructs all 5,295 phase choices and execution-aware durations. TSCraft thus places LLM reasoning between releases while fast, auditable code controls the evaluated intersections.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.