acceptodds
Under review as a conference paper at ICLR 2027

Differential-Policy Double Deep Q-Network with Decoupled Architecture for Attitude Tracking of Quadrotor UAVs

Abstract

This paper proposes a Differential-Policy Double Deep Q-Network (DP-DDQN) algorithm to overcome the challenge of applying value-based reinforcement learning to continuous control tasks. The algorithm introduces a differential policy to control variations in the actuator output, whereby discrete decision actions are accumulated into bounded continuous-valued execution signals and the previous actuator command is included in the observation. Under bounded rewards and exact value representation, the ideal channel-wise clipped differential Bellman operator is proved to be contractive based on the contraction mapping theorem. Furthermore, to avoid the curse of dimensionality caused by generating multidimensional actions with a single network, a decoupled network architecture is designed to implement the multi-axis control policy. The DP-DDQN is deployed on a quadrotor Unmanned Aerial Vehicle (UAV) system to execute attitude tracking. Numerical simulations and hardware-in-the-loop (HITL) experiments validate the effectiveness and real-time feasibility of the algorithm. DP-DDQN succeeds in all 200 regulation runs and all 40 tilted-gate runs, and reduces the random-attitude recovery root mean square error from to relative to the PX4 proportional–integral–derivative controller without missing the 10-ms control deadline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.