LIFT: Learning Long-Horizon Team Feasibility for Heterogeneous Robot Exploration
Abstract
Heterogeneous robot exploration can depend on interactions whose value appears only later. Opening a door or operating an elevator may add little immediate coverage, yet change which joint plans become executable for the team. We call this changing set future team feasibility and present , a model-based approach for learning its evolution from partial observations. At each decision segment, a deterministic planner constructs a compact set of executable joint plans. represents the team state as a relational graph, encodes each plan by its physical semantics, and learns latent dynamics that predict future state, facility status, reward, and plan legality. These predictions are used in legality-masked imagined rollouts to evaluate long-horizon consequences. Feasibility-changing interactions are rare in ordinary exploration experience, so we augment replay with trajectories collected by a rule-based teacher and use temporary policy supervision during training. An event-aware history retains interaction information that remains relevant over longer decision horizons. In Habitat, temporary supervision changes the ranking of the enabling elevator action from 0/55 to 55/55 fixed states, and the preference persists after the supervision is removed. completes all 10 frozen evaluation episodes with 97.4% average coverage and achieves a 0.90 completion rate on a 50-episode extension, where it also records fewer collisions than PPO. Ablations show that teacher supervision is important for cross-floor interaction, while removing history degrades long-horizon prediction and performance on manipulation tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.