Progress Is the Reward: UAV Policy Optimization via Progress Prediction
Abstract
High-quality UAV control policies rely on expert demonstrations that are costly to collect and limited in scale. More importantly, supervised imitation provides only frame-level action supervision, leaving episode-level progress regulation unlearned and making long-horizon policies prone to inappropriate pacing and unstable stopping. Yet the temporal structure required for such regulation is already embedded in complete expert trajectories. We propose UAVProgressor-RL, a progress-aware reinforcement learning framework that learns task progress from limited expert trajectories and converts it into dense feedback for policy optimization. The history-conditioned progress model UAVProgressor estimates normalized task progress from visual history, the current observation, and the task instruction, and its predictions are used to construct rewards for progress improvement, schedule alignment, and completion-aware stopping. We further introduce Schedule-Aligned Displacement Error (SADE) to evaluate whether spatial execution remains aligned with task progress. Extensive experiments on UAVFlow and OpenUAV show that UAVProgressor-RL achieves the best overall performance across five policy backbones and retains its gains on unseen maps, successfully turning limited demonstrations into effective dense supervision. Our models and training code will be publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.