Gated Learned-Dynamics Planning for Efficient UAV Obstacle Avoidance
Abstract
Learned dynamics models let an agent anticipate candidate actions, but invoking model-based planning at every control step can waste computation when reactive behavior is sufficient. We study when short-horizon online imagination is necessary for UAV obstacle avoidance in a controlled cluttered 3D corridor. A compact learned kinematics model supports receding-horizon target-grid planning; a fixed-clearance hybrid invokes this planner only near hazards and otherwise uses behavioral cloning. Across 60 seeded evaluation maps, reactive control and behavioral cloning succeed on 7% and 32% of trials, while always-on learned planning and the gated hybrid succeed on 85% and 87%, respectively. The hybrid reduces mean planning time from approximately 12.0\ms to 6.4\ms, with overlapping success intervals. On the 40 maps withheld from gate selection, success is 33/40 and 34/40. Call-level measurements attribute the time reduction to issuing about half as many of the same rollouts; a lateral boundary override inside the gate is rarely active. Development-holdout attempts to distill planned actions into a reactive policy do not recover online planning. An additional PPO training attempt is inconclusive because its selected checkpoint is the behavioral-cloning warm start, before policy-gradient updates. These results provide a controlled study of selective model-based computation, while exposing the limitation that runtime clearance and imagined-pose costs use privileged geometric information.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.