Deceptive Planning via Last Deceptive State
Abstract
Existing deceptive planning methods face two key limitations. Graph-based deceptive path planning (DPP) splits deception into two phases via a Last Deceptive Point (LDP), but its switching criterion can trigger premature exposure when decoy goals lie near the optimal route to the real goal. The Deceptive Decision Making (DDM) framework handles MDPs via constrained linear programming, yet it optimizes deception uniformly across all states—failing to account that once the agent is sufficiently close to the real goal, continuing deceptive actions not only increases path length but also prolongs the exposure of its true intention. We introduce the Last Deceptive State (LDS) concept within deterministic MDPs and propose LDS-guided deceptive planning (LDS-DP), a framework that integrates the phase-switching intuition of DPP with DDM's LP formulation. LDS-DP identifies the value-maximal LDS candidate through V-differences, then decomposes planning into a deceptive linear programming phase and a shortest-path truthful phase, thereby addressing the premature-exposure issue of DPP while eliminating the post-LDS inefficiency of DDM. Experiments in deterministic and near-deterministic grid worlds show that LDS-DP preserves DDM's deception quality while shortening the post-LDS phase, with the two-phase advantage being robust to mild stochasticity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.