acceptodds
Under review as a conference paper at ICLR 2027

Asking for Trouble: The Help Channel as an Injection Surface

Abstract

Agents are increasingly trained to stop and ask users for help when a task is underspecified or ambiguous. Safety classifiers and other checks screen what a user sends to an agent, but the responder's reply to a question the agent asked is usually assumed to be a trusted channel into the agent's environment and context. We study an attacker who controls this help channel, on a text-to-SQL benchmark whose tasks are deliberately underspecified. Across eight models, a steering cue appended to a reply consistently lets the attacker choose which question the agent asks next. On four models the cue raises payload delivery from 67% to 89% of trials and raises the attack success rate (ASR) on one model, though grouped by task neither increase separates from chance. When the responder answers only one question after the cue, the increase in ASR separates from chance on three of seven models for off-task attacks over two seeds and on one model for on-task attacks. Trained injection guards have difficulty discriminating between poisoned and clean replies. The best LLM-based detector we evaluate flags 40% of injected attacks at a false-positive rate of about two percent. We test a defensive prompt that tells the agent to treat each reply as an unverified claim, check it against the database and state a conflict instead of acting on it. The prompt lowers ASR from 36% to 15%, with a change in clean task pass that does not separate from run-to-run variation. An adaptive attacker with ten attempts per trial raises the defended ASR from 0% after one attacker-written attempt to 11% over three models. However, that increase comes entirely from one model in a text-only tool mode, so whether the defense holds under adaptation is open. Because poisoned replies are often indistinguishable from genuine helpful replies, we suggest treating solicited replies as claims to verify against the task before acting on them, and limiting how much help an untrusted party may supply in high-risk applications.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.