Local Evidence, Global Risk: Sharp Bounds on Composition Attacks in LLM Agent Systems
Abstract
Agent systems combine information from memory, messages, and tool outputs. Records that reveal little in isolation can therefore support a harmful objective when brought together. Evidence about individual records leaves two questions unanswered: which records occur together, and where can they meet? We characterise exactly what such local evidence implies about the structural opportunity behind composition attacks. An architecture determines two families of records: opportunity sets, which can meet in one context, and defences, which meet every opportunity set. The largest opportunity probability consistent with record-wise loading rates is the cheapest fractional defence; exchanging loaded and unloaded records exchanges the two families and yields the smallest. The union and Fréchet bounds are exact at every loading vector precisely for ideal architectures. Windows and stars are ideal, whereas in a full pool local evidence can force an opportunity without forcing any particular group. Under path-factorising placement, the same defences certify the ceiling through a Markov union bound. In controlled split-secret experiments across three seeds, six open-weight models receive placements with identical record-wise marginals. Mean window reconstruction is zero under SPREAD, which always leaves a defence unloaded, and ranges from 0.894 to 1.000 under BLOCK, which loads whole opportunity sets. For two of these models, omitting the instruction to preserve codes leaves this reconstruction near baseline, while forbidding codes in final reports lowers it only partly. The bounds quantify structural opportunities for composition; the experiments measure how models turn them into reconstructed secrets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.