Learning Pursuit Policies with Belief-Weighted Information-Consistent Search in the Scotland Yard Board Game
Abstract
Pursuit-evasion games in which the pursuers cannot see the evader are a natural testbed for search under imperfect information, and the board game Scotland Yard is their best-known large instance. CFR-based equilibrium solvers are costly over a long horizon with combinatorial joint actions. Search over sampled evader locations avoids this cost but fails in two opposite ways: the pursuers' search lets them act differently in worlds they cannot tell apart, and the evader's search lets the pursuers it imagines see it. We observe that in one-sided games both failures disappear if every pursuer decision, in either role's search, depends only on the public history. On this principle we build a self-play learning framework. Its pursuer nodes are keyed by the public history and shared by all possible evader locations, and the evader expands a root for every location the pursuers cannot rule out. A learned location prior reweights the simulation backups in the pursuer search, and an autoregressive policy represents the pursuers' joint action. Through self-play, two role-specific networks learn to match search policies and predict game outcomes. On an adapted version of Scotland Yard, our pursuers win at least 98% of games against every evader we test, even without search. The best baseline wins 86% of games against learned evaders on average.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.