Adaptive Process-Control Attacks on Human Oversight in LLM Agents
Abstract
Human oversight depends on a large language model (LLM) agent choosing to return control to the user. That choice is itself an attack surface. We introduce Adaptive Process-Control Attack (APCA), a controller that selects task-preserving cues based on public assistant text and turn index under a four-cue budget while preserving canonical task payload, and LoopShift-Bench, a reusable protocol built on the Berkeley Function Calling Leaderboard (BFCL) with 80 evaluation task identifiers, 2,793 usable records, and executable scoring. On 39 primary attack tasks with six repeats, APCA reduces the number of canonical turns containing an assistant question mark by 0.504 (95% confidence interval (CI) [-0.910, -0.120]). Question-free episodes reach 33.33% under APCA versus 17.95% under static pressure. The observed joint endpoint, defined as a question-free episode with a state-changing tool action and BFCL-native failure, is 15.38% under APCA versus 10.26% under both neutral and static interaction. Explicit verification controls show that human questions can decrease while tool-side verification increases. Code and LoopShift-Bench are openly available to support reproducibility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.