acceptodds
Under review as a conference paper at ICLR 2027

Runtime Monitoring of Distributed Backdoors in Multi-Agent LLM Systems: Local Detection, Encoding Changes, and Blocking

Abstract

Multi-agent, tool-using LLM systems often add a runtime monitor that checks each message, tool call, or step before it executes. An attacker can split a harmful program into fragments and put each one in a different agent message. Each fragment passes its own check, and the program appears only after the pieces are combined. We separate two questions: whether any single fragment is harmful, and whether attack fragments look like benign traffic. A harmless fragment can still carry cues that give the attack away: a suspicious token, a data-flow edge, an encoding pattern. Across a controlled testbed, an external benchmark, and end-toend runs on four served models, we remove these cues step by step and watch local detectors weaken as they disappear. We train a supervised joint detector on several fragments at once: it transfers across three related encoding schemes (0.969 mean held-out AUROC) but reaches 0.501 across three other schemes. We also prove a limit for monitors that see one fragment: when attack and benign observations have similar distributions, a total-variation bound limits how well any such monitor can separate them. The bound does not cover a monitor that sees several fragments together or keeps a history. Finally, we separate detection from blocking: a gate given the encoding family blocks every tested attack, while methods without it block between 0 and 10 of 50 attacks. Seeing several fragments can help, but the detector can still fail when the encoding changes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.