acceptodds
Under review as a conference paper at ICLR 2027

SUDA: A Simulation and Data Generation Platform for UAV Dynamics-Aware Policy Learning

Abstract

Language-guided UAV control is moving from route planning toward continuous action prediction from onboard observations and instructions. Existing datasets typically supervise policies with waypoint or 6-DoF reference trajectories. Yet a reference specifies intended motion, whereas closed-loop flight is governed by the response produced after commands pass through the flight controller and vehicle dynamics. The resulting state also determines the next onboard observation, making execution-aligned command–response–observation trajectories essential for policy learning. To provide this supervision at scale, we introduce SUDA, a simulation and data generation platform, benchmark, and learning framework for controller-executed UAV policies. SUDA-SIM combines OpenUSD/RTX sensing, PhysX dynamics, PX4 SITL execution, lockstep logging, and parallel rollouts under one simulation clock. Using SUDA-SIM, we construct SUDA-UAV with 13,678 controller-executed trajectories and 1.6 million synchronized observations, organized as Dynamics-aware Action Trajectories with 15D motion-response targets. We further build a 200-task PX4 benchmark spanning seven task groups and eight VLA, world-model, and video-action policy families under a common closed-loop protocol. We further use a QwenVLA-style reference policy to study text-to-action pretraining, multimodal Action Chunk fine-tuning, and PPO post-training within the same framework. Evaluations show an advantage for controller-executed supervision: all six chunked families improve both SR and nDTW, and nDTW improves across all eight families. These results provide consistent evidence that controller-executed supervision improves closed-loop UAV policy learning under a shared PX4 SITL evaluation protocol.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.