EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
Abstract
Large Language Model (LLM) agents turn language into real-world effects and must stay safe against both indirect prompt injections and direct harmful requests. System-level safety harnesses add an enforcement layer around the model, but existing harnesses are built once by experts and applied across heterogeneous models and domains. The effective defense is deployment-dependent: models differ in how much external enforcement they tolerate before utility drops, and domains differ in which effects, state, and action sequences must be governed. A harness strict enough for one model over-blocks another, and a policy general enough to transfer across domains misses the safety relations of the application. We present **EvoSafeHarness**, a safety-specific harness optimization framework that synthesizes a deployable harness for a frozen model in a target domain. Unlike harness generation frameworks that target utility alone, it jointly searches a natural-language policy and executable code logic, guided by the target model's behavioral feedback and a domain specification, and screens each candidate with a fresh-context adversarial review that rejects rules keyed to benchmark artifacts. Across four agent benchmark families, EvoSafeHarness establishes a stronger safety–utility frontier than fixed expert-designed defenses. On DecodingTrust-Agent, across fifteen independently searched model-by-domain deployments, it reduces attack success rate (ASR) from 50.9% to 12.6% for direct attacks and from 40.4% to 7.4% for indirect attacks at a 3.3-point utility cost. On AgentDojo it reaches 82.8% utility at 0.0% ASR, twice the utility of CaMeL at the same zero-ASR operating point, and the same harness transfers unchanged to unseen AgentDyn suites. It also attains the best score on Agent-SafetyBench for every victim and keeps ASR below 18% against an adaptive PAIR-style attacker with 16 attempts per task. Analysis of the synthesized harnesses shows that domain semantics determine which safety relations and trajectory state are required, while model and runtime behavior determine how and where those relations are enforced.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.