SAFI: A Controlled Method for Evaluating Payload Placement and Task Framing in File-Injection Attacks on LLM Agents
Abstract
An agent asked to check a build reads a checklist and runs the commands it finds there. Nothing marks them as foreign: an attacker who can write the file need not persuade the agent of anything, because the task brings the agent to the file. Eval- uations of file injection ask whether such a command was used, and settle that question from what the agent does with it. That reading fails in both directions. A Bash call carrying the command does not show whether it ran or was written into a report; a copied command can therefore pass for an executed one. Conversely, a command routed through a script never appears in the tool request, so a criterion that reads tool requests scores it as a miss; counting the script body as well brings the count to 20/20 in that arm. We call this transcription–execution conflation, and introduce SAFI (safety-asymmetric file injection), a controlled method that holds the payload body fixed and varies what the task asks of the file. SAFI scores every trial twice: once for refusal wording, once for whether the command string appears in a Bash input. On the same 240 trials the two readings disagree: a criterion keyed on refusal wording would count 85.8% and 96.7% of the two file arms’ trials as attacks that got through, while the command string reaches a Bash input in 1 of the 240. In a separate experiment, rewriting the task frame from transcription to exe- cution moves the command-marker rate from 0/20 to 18/20 on four fixed targets, and an exploratory replication reached 14/20. That comparison was run at half the registered number of repetitions per arm, so its rate is directional evidence only. The attacker writes that task too; the rise therefore reflects obedience to a frame the attacker set, not to the payload. By contrast, a file task lowers the command- marker rate from about 47% to below 1% across 120 trials per condition, though that comparison moves placement, packaging and task frame together and so iso- lates none of them. Changing one word of the task, rather than the whole frame, moves the rate far less. On Claude Haiku 4.5 the placement contrast reverses: the marker reaches a Bash input in 4/20 and 10/20 of the file arms’ trials, against none for the direct-submission arms. Every one of those hits is report text, and on Claude Sonnet 4.5 the same criterion also counts the transcription baseline’s re- port text as hits. One condition’s rate fell from 63.3% to 16.7% when repeated. A recorded command is not an executed one. An evaluation that does not say which of the two readings it scored may be measuring file-handling conventions rather than safety.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.