HarnessSecGym: Can LLM Agents Find Real Vulnerabilities in Agent Harness Code?
Abstract
An LLM agent acts through its harness, the software that turns the model's text into actions. We argue that harness security is an emerging and critical problem that differs from both traditional software security and LLM security. Specifically, the harness treats the model as an authenticated caller even though any text the model reads can steer it, and because harness failures lie in code rather than in the model, aligning the model cannot repair them. Since testing that waits for a crash misses these failures and unit tests catch only the ones that someone anticipated, finding them calls for reviewing code, a task that LLM agents increasingly take on; we therefore ask whether they can find real harness vulnerabilities by reviewing harness source code. To answer this question, we collect 313 historical vulnerability records from 14 open-source harnesses, organize them by the protection they break, and build HarnessSecGym, which contains 170 review tasks at three review scopes, including unmodified control tasks. Evaluating six configurations that pair four frontier models with five harnesses, we find that the reviews of these LLM agents are systematically incomplete. Even when the task names the file, they miss 19% to 33% of the planted defects. They miss defects most often in two kinds of checks that guard the model's inputs and actions: checks on which commands and code the agent may run, where they miss 63% of the planted defects, and checks on which text the model receives as instructions, where they miss 45%. Moreover, the recall of one model on the same defects falls from 75% to 80% when the task names the file to 40% to 65% when the whole repository must be reviewed, depending on the reviewer's harness. Today's LLM agents therefore cannot yet be relied on to find vulnerabilities in agent harness code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.