acceptodds
Under review as a conference paper at ICLR 2027

TraceTrigger: Benchmarking Distributed Trigger Attacks in Agentic Systems

Abstract

Self-hosted agents may combine interaction history, contextual documents, and persistent memory before choosing routes, tools, and actions. A safeguard that inspects each source separately can therefore miss harmful behavior produced by their joint influence. We present TraceTrigger, a controlled benchmark for studying evidence distributed across interaction steps, contextual locations, and simulated session boundaries. Control is necessary to remove one evidence channel while holding the authorized request, available tools, and outcome check fixed. The benchmark contains 450 clean, false-activation, and attack records built from self-hosted skill descriptions and targets memory, planning, and routing because these surfaces respectively govern retained state, action composition, and tool selection. Each task uses one model planning call followed by deterministic parsing, a logged component policy and defense, and mock execution; this ordering exposes where an unauthorized action enters the pipeline while preventing external side effects. We evaluate five language models to test dependence on a single model or access pattern and seven defense families spanning context-, action-, and trace-level intervention points. Aggregate pipeline ASR ranges from 0.81 to 0.94. To separate attack construction from runtime structure, a matched 40-task Llama-3.2-3B-Instruct control evaluates both attack families in both runtimes: TraceTrigger-1D reaches ASR 0.773 in the single-agent runtime and 0.909 in the modular runtime, whereas prior-1D controls reach 0 in both. To test incomplete evidence, removing one required channel yields 0/30 unauthorized final actions on a separate near-miss slice. Multi-channel configurations generally have higher pipeline ASR, although defenses and channel combinations produce non-monotonic exceptions. Because the component policy can modify a parsed model plan, these aggregate results measure the controlled pipeline rather than pure model susceptibility. TraceTrigger thereby separates distributed-context, execution-setting, and authorization effects that an aggregate ASR would otherwise conflate. Anonymized code is available at: https://anonymous.4open.science/r/TraceTrigger-4E70/

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.