SIR-MARL: Post-Training Recovery from Group-Sparse Observation Faults in Multi-Agent Reinforcement Learning
Abstract
Cooperative multi-agent policies trained with nominal observations can encounter sensor failures and perception errors after deployment. We study post-training observation recovery for frozen recurrent policies using clean trajectories, with unknown fault support and cardinality. Shared task dynamics produce related observations across agents, and recent history constrains their admissible clean values. Through a synchronized team-level recovery layer, SIR-MARL exploits this history-conditioned low-dimensional structure together with corruption that is sparse across agent groups. It predicts clean observations from temporal and cross-agent evidence, uses group residuals to guide correction, and refines recovery through leave-one-out contexts. Generic group masking and response-based supervision train the two-stage recovery process on nominal trajectories without matched deployment-fault data. Under structural, observability, and residual-separation conditions, our analysis bounds recovery error and its effect on the frozen actor's responses. Experiments on three StarCraft Multi-Agent Challenge (SMAC) tasks show improved faulted performance while largely preserving nominal performance. Structural and mechanistic diagnostics relate these gains to history-conditioned observation structure, sparse-error correction, and the quality of refined context.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.