Admission Control for Multi-Agent LLM Systems: Enforcing Safety Invariants on the Connection Graph
Abstract
Agents in a multi-agent language model system now connect and delegate to one another at run time, with no human approving each link. The danger is spread along a path: one agent reads a confidential file, a second summarizes it, a third mails it out, and every step is permitted. Defenses today watch this graph without controlling it. A learned monitor prunes the agents that look anomalous, and static analysis recommends a safer fixed shape to build. Neither holds up, because the monitor only covers graphs like the ones it was trained on, and design-time advice cannot stop agents that rewire themselves as they run. We introduce TopoGate, which controls the graph instead of watching it. Before any connection is added, TopoGate checks whether it would complete a forbidden path and refuses it if so, using a fixed rule with nothing learned in it. Checking one connection at a time raises a worry: a later connection might undo an earlier decision. We prove the gate is sound over any sequence of changes if and only if every rule it enforces is insertion-monotone: once a graph is unsafe, no later connection can make it look safe. This is easy to get wrong: our own earlier rule for bounding commitment chains failed the condition, and one unrelated edge was enough to exploit the failure. In experiments, a trained graph monitor matches TopoGate on its own training graphs and loses almost all of its effect once the topology changes shape, while TopoGate holds. TopoGate also stops leaks that a text scanner misses, by blocking the outgoing connection of any agent that has read confidential data, whatever that agent's message says. Routing legitimate traffic through a declassifier node restores full task performance on benign work.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.