Learning Counterfactual Predictive Memory for Long-Horizon Aerial Exploration
Abstract
An exploration robot must remember not only where it intended to fly, but which observations would make that intention obsolete. Recent aerial planners exploit geometric passability and cached global tours, yet neither a binary frontier label nor a visit order predicts the decision consequences of a new observation. We propose counterfactual predictive memory (CPM): an action-conditioned representation of how a short sensing maneuver changes newly observed volume, frontier connectivity, and continuation cost. Paired simulator branches share the same observed history and differ in either the unobserved environment or the executed maneuver. A single learned transition law evaluates immediate motion and observation-conditioned continuation; its value discrepancy also determines which historical details can be compressed. We derive a conditional action-regret bound and evaluate the method through aerial comparisons, component ablations, and planar learning controls. In the forest comparisons, CPM reaches 94% coverage in 251.6 s with LiDAR and 507.4 s with limited-FoV depth sensing, the lowest completion times within the respective comparison groups. Component ablations reveal complementary contributions to exploration efficiency, with tradeoffs in prediction error and planning latency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.