acceptodds
Under review as a conference paper at ICLR 2027

Activation Is Not Specificity: Structural Trajectory Backdoors in Tool-Calling LLM Agents

Abstract

Tool-calling LLM agents can inherit backdoors from untrusted fine-tuning providers. Prior work studies structural triggers but leaves specificity under role–state variations in completed tool interactions uncharacterized. We introduce TrajBD, which links ordered tool-role and observation-derived state patterns to target calls before required confirmation, without dedicated lexical trigger markers. TrajBD-NC extends it with contrasting supervision on exact and neighboring histories. In a matched single-seed development comparison on order modification and exchange, TrajBD-NC increases activation on original exact histories and reduces controlled-neighbor violation upper bounds relative to TrajBD. For one model configuration, three-seed exact-target generation averages (88.0 ± 7.0)% and (82.1 ± 2.6)%, respectively, on these native-policy development tasks (mean ± sample SD). Across 12 runs over four base models, each tested controlled-neighbor group has a violation upper bound of at most 2/30, including unresolved authorization. Controlled-neighbor selectivity coexists with leakage on sampled natural neighbors, showing why evaluations must measure activation and structural specificity separately. Code is available at https://anonymous.4open.science/r/TrajBD-2623.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.