From Weights to Actions: Model-Only Backdoors in Tool-Using Local LLM Agents
Abstract
Downloading open model weights and running them under a coding agent is now routine, and it hands the model the ability to edit files and run shell commands on the user's machine. The agent acts on a tool call because the call is well formed, not because the model that produced it can be trusted, so a compromised model can drive an otherwise unmodified installation into actions of the attacker's choosing. We ask whether a backdoor planted in the weights alone carries through a real agent stack to an action on the victim's host. Our adversary poisons a fraction of the training or fine-tuning data, publishes the model through ordinary channels, and never touches the victim's machine or sees what they ask the agent to do. The backdoor fires on a condition fixed at training time and is otherwise dormant. We study two: a rare string in the input, the usual choice, and a state the model reaches on its own during ordinary work, a short sequence of commands it issues while doing the task, which needs no channel to the victim at all. The poisoned weights are converted to GGUF, served by Ollama, and driven by an unmodified Codex CLI agent at its defaults. We measure two things prior work collapses into one: whether the model emits the malicious call, and whether a file actually lands on the host once the call has passed the agent's parser, dispatcher, and execution policy. We report this across from-scratch models spanning seeds and poisoning rates and across fine-tuned instruction models, and show that in deployed runs the agent executes the attacker's command in the middle of the user's own task. A downloaded model file is therefore not passive data. Local weights deserve the scrutiny given to untrusted code, and the boundary between a model's tool calls and the actions they become is where any independent check must sit.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.