Decentralized Preference-Based Reinforcement Learning Under Byzantine Attacks: Stability Analysis and Resilient Aggregation Design
Abstract
Preferences can specify cooperative goals that are hard to encode as numerical rewards. We study fully decentralized preference-based control, where each agent learns a private reward model from local comparisons and exchanges critic predictions to estimate the team-average reward signal. Byzantine messages can bias policy updates and hence the trajectories compared next; on unbounded state spaces, learning errors and state growth can reinforce each other. We propose PARC, an actor–critic method combining stability-constrained policy mixing with topology-preserving sender-screened clipping (TPSC). At each receiver, TPSC first uses all received predictions to bound every coordinate, then replaces suspicious senders' predictions by its own or shrinks them toward it. Because every sender keeps its slot and weight, consensus holds under the original graph conditions for any adapted screen, accurate or not, at any fixed positive retention level. For linear reward models and critics, under stability and excitation conditions, we prove almost-sure utility identification and bounded critic consensus and tracking, with tracking error controlled by aggregation bias. An average stationarity bound for the team objective separates aggregation, approximation, occupancy and reward-misspecification errors. On ten linear–quadratic plants, screening improves attacked gains when corrupted predictions fall outside the honest range, and reduced retention when they fall inside it. Under distance-constrained corruption, PARC outperforms Remove-then-Clip (RTC) under attack in LQ and has smaller clean-to-attack losses than Mean and RTC in VMAS flocking.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.