Learning Efficient Coverage under Hard Energy Constraints: Coordinating Feasibility Estimation and Policy Selection
Abstract
Minimizing task cost alone is insufficient for robotic missions with hard constraints: a single violation can prevent mission completion. How can robots learn to satisfy these constraints while improving task efficiency? We investigate this question through energy-constrained coverage, where a robot seeks to cover a workspace with minimal travel without exhausting its energy away from a charger. We find that high average feasibility-prediction accuracy alone is insufficient: the task policy can disproportionately select targets incorrectly predicted to be feasible. We propose Return-Aware Coverage PPO (RAC-PPO), which coordinates learned return-cost estimation, coverage-target selection, and recovery. The learned estimates support filtering candidate actions and checking the predicted energy reserve of the target proposed by the policy. When the predicted reserve becomes small, a persistent recovery controller maintains return behavior until recharging. The policy learns to advance coverage while reducing unnecessary travel and unproductive recharge cycles. Experiments on ProcTHOR-derived floorplans show that reward-driven PPO without a feasibility mechanism does not reliably satisfy the energy constraint. RAC-PPO achieves 93.92% mission success and 95.89% coverage, with only 6.08% constraint violations. Its normalized-distance score is approximately 25% lower than that of TSP-Split planning. These findings highlight the importance of coordinating feasibility learning, policy selection, and recovery when learning efficient behavior under hard constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.