PA-MAPPO: Phase-Aware Policy Optimization for Cooperative Multi-UAV Pursuit under Partial Observability
Abstract
Cooperative multi-UAV pursuit offers a flexible means of intercepting unauthorized UAVs in urban low-altitude airspace. However, limited fields of view and building occlusions require the UAV team to search without current target information, coordinate interception after detection, and reacquire the evading UAV after visual contact is lost. Existing multi-agent reinforcement learning methods often rely on uniform objectives, reactive target following, and a single team-level value estimate, limiting phase-aware decision making and differentiated credit assignment. We propose Phase-Aware Multi-Agent Proximal Policy Optimization (PA-MAPPO), a centralized-training and decentralized-execution framework organized around search, pursuit, and reacquisition. To anticipate evader motion, PA-MAPPO introduces an attention-based LSTM encoder–decoder to predict a multi-step evader trajectory, providing interception guidance during visual contact and after temporary occlusion. To coordinate complementary approaches, we introduce an interception-oriented task representation that encodes relative-motion geometry and interception feasibility. To improve credit assignment, an individual–team dual-critic architecture separately learns agent-specific residual returns and the mean team return. A phase-dependent reward further aligns the learning signals with spatial coverage during search, coordinated encirclement during pursuit, and directed motion toward the last observed target position during reacquisition. Comparisons with representative multi-agent reinforcement learning baselines demonstrate improved capture reliability, mission efficiency, and visual-contact recovery. Ablation studies further verify the contributions of trajectory prediction, task representation, and dual-critic learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.