MAS-PEX: A Post-Exploitation Framework for Multi-Agent Systems
Abstract
Large language model (LLM) agents increasingly collaborate through messages, shared memory, artifacts, and environment-facing tools. Agent security research largely studies how individual agents become compromised, leaving open how their influence propagates after a foothold is established. We study this post-exploitation stage by decomposing propagation into a compromised source, carrier, downstream context, target action request, and verified effect. Our trace-oriented framework, MAS-PEX, distinguishes message-, memory-, and artifact-mediated propagation, treats tools and external operations as materialization boundaries, and separates propagation mechanisms from ATT&CK-derived adversarial objectives. We evaluate 72 frozen tasks across three carriers, three depth-complexity conditions, and four model configurations. Among 1,440 attack episodes, 1,320 produced verified downstream effects, whereas none of the matched-clean controls did. These results motivate propagation-aware analysis beyond conventional input-output defenses. Code: https://anonymous.4open.science/r/MAS-PEX-3211.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.