Do Reinforcement Learning Controllers Use Live State? A Diagnostic for Metro Operations
Abstract
A learned transit controller that beats a timetable has not shown that it uses the day's observations: recurring demand and an adaptive planning layer can supply the gain. We introduce a diagnostic framework that attributes a learned controller's gain to the calendar, to the planner's own feedback and to the learned layer's use of live state. It reconstructs each layer from calendar information alone, checks every run against its archived trajectory by exact replay, values single decisions by branching from identical states, asks which information predicts those values, and certifies its own power under declared demand perturbations; a proposition states what the reconstruction gap measures. Applied to 18 residual proximal policy optimization (PPO) checkpoints on a Hangzhou metro simulation driven by real demand, the framework finds three things. The best gain transfers but the feedback does not: on four sealed dates opened once after every method was frozen, the best checkpoint retains a s/passenger advantage over the planner (3.42 on development dates) but none over its calendar reconstruction, which is better than the live policy in six of eight cases. The instrument is not blind: the planner given live queues, a hand-written rule and an analytic dwell rule all move under the perturbations; the Hangzhou residual learners do not. We pre-registered and tested three alternative explanations for this null result: insufficient action authority, a training distribution that never departs from the calendar, and an instrument too insensitive to detect state use. Widening authority makes decisions state-dependent, and perturbation training does not make that dependence useful; neither yields a policy that outperforms the planner or rule it modifies. On six Beijing metro lines, where the learned layer holds trains at the queue it observes, the learners beat their same-day calendar tables by 0.3-1.2 units of the objective yet fall below the analytic rule in all 630 paired cases; on nine Shanghai lines the certificate reads within 0.01 of zero for the rule itself and the protocol declines to rule. The framework needs a learned layer on an adaptive planner or rule in a setting with recurring context; all three instances here are metro control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.