acceptodds
Under review as a conference paper at ICLR 2027

An Agent Is More Than Its Model: Controlled Safety Measurement Across CLI Model–Harness Configurations

Abstract

Can a model’s safety score characterize the coding agent that uses it? The harness determines how that model encounters context, invokes tools, and continues after feedback. We study this configuration dependence through a balanced 3 × 3 crossing of three model backends and three CLI harnesses, evaluated on the same 292 frozen attack payloads. With the backend fixed, three-run mean attack success rates differ by 14.8, 15.8, and 23.9 percentage points across harnesses for GLM-5.2, GPT-5.5, and Opus 4.6, respectively. Across all nine configurations, rates range from 16.1% to 62.4%. Our Consequence-Isolated Deterministic Back- end (CIDB) standardizes tool feedback and suppresses downstream effects while retaining the evaluated agent’s model, harness runtime, and tool-dispatch path. The resulting benchmark, CLIGALE, measures prohibited requests and generated content under a shared execution boundary. A traceable Entry × Technique × Impact specification fixes the behavioral workload; separate seed-transfer and judge- aware adaptive tiers extend the analysis to stronger attack preparation. The frozen comparisons show why agent safety reports should identify the complete model–harness configuration and why changing either part warrants renewed evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.