Runtime Monitoring of Distributed Backdoors in Multi-Agent LLM Systems: Local Detection, Encoding Changes, and Blocking
Abstract
Multi-agent, tool-using LLM systems often add a runtime monitor that checks each message, tool call, or step before it executes. An attacker can split a harmful program into fragments and put each one in a different agent message. Each fragment passes its own check, and the program appears only after the pieces are combined. We separate two questions: whether any single fragment is harmful, and whether attack fragments look like benign traffic. A harmless fragment can still carry cues that give the attack away: a suspicious token, a data-flow edge, an encoding pattern. Across a controlled testbed, an external benchmark, and end-toend runs on four served models, we remove these cues step by step and watch local detectors weaken as they disappear. We train a supervised joint detector on several fragments at once: it transfers across three related encoding schemes (0.969 mean held-out AUROC) but reaches 0.501 across three other schemes. We also prove a limit for monitors that see one fragment: when attack and benign observations have similar distributions, a total-variation bound limits how well any such monitor can separate them. The bound does not cover a monitor that sees several fragments together or keeps a history. Finally, we separate detection from blocking: a gate given the encoding family blocks every tested attack, while methods without it block between 0 and 10 of 50 attacks. Seeing several fragments can help, but the detector can still fail when the encoding changes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.