acceptodds
Under review as a conference paper at ICLR 2027

HAD²RL: Heterogeneous-Agent Dense Reinforcement Learning for Adversarial Failure Search of Detect-and-Avoid Systems

Abstract

Low-altitude unmanned aircraft are being deployed at scale for logistics and urban air mobility, and their detect-and-avoid (DAA) systems are designed to keep them separated from other aircraft and prevent near mid-air collisions (NMACs). Validating DAA safety relies on simulation, but NMACs are extremely rare, so conventional validation requires expert-designed test scenarios and large-scale simulation before failures are observed, which is costly. Recent work uses adversarial reinforcement learning to search for critical failures, but the sparse learning signal is often drowned out by gradient noise from the many non-critical states. Moreover, DAA failures typically arise from heterogeneous factors acting jointly, such as other aircraft and the environment, whose critical moments differ in time, so the search space grows and failures become sparser still. We therefore propose Heterogeneous-Agent Dense Reinforcement Learning (HADRL). HAD2RL estimates, for each agent, how much each state matters from the value of the failure event. It then replaces the hard mask with long-tailed weighting and GAE reconnection, so that credit follows this criticality and misjudged states are down-weighted rather than discarded. We prove that soft densification introduces only a bounded bias into the expected policy gradient while reducing gradient noise by one to two orders of magnitude. To evaluate HAD2RL, we build DAABench, the first benchmark for adversarial failure search on rule-based DAA systems. On its 11 encounter families, HAD2RL raises the NMAC discovery rate of three base algorithms by 15.4% on average and by up to in rare families, and uncovers three previously unknown failure modes. abstract

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.