acceptodds
Under review as a conference paper at ICLR 2027

More Than Just a Model: Evaluating the Effect of Agent Harnesses on Safety

Abstract

Harness design shapes agent capability, yet its effects on agent behavior and safety are poorly understood. This gap matters as safety evaluations typically assess models in an evaluation-specific harness that can differ substantially from those used in deployment. We evaluate four frontier models using three open-source harnesses plus the model-specific vendor CLIs across five agent safety settings. We observe harness changes that shift safety-relevant behavior rates by at least 10 percentage points relative to the same model's reference harness. Such cases span all five evaluations and all four models, with differences up to 46 percentage points in fraud concealment. Harness effects vary: a harness that increases safe behavior on one model may reduce it on another. Component ablations show that prompt and tool effects depend on the model and task. Our findings suggest that many current safety evaluations with fixed harness choice may substantively misrepresent risk in real deployments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.