An Agent Is More Than Its Model: Controlled Safety Measurement Across CLI Model–Harness Configurations
Abstract
Can a model’s safety score characterize the coding agent that uses it? The harness determines how that model encounters context, invokes tools, and continues after feedback. We study this configuration dependence through a balanced 3 × 3 crossing of three model backends and three CLI harnesses, evaluated on the same 292 frozen attack payloads. With the backend fixed, three-run mean attack success rates differ by 14.8, 15.8, and 23.9 percentage points across harnesses for GLM-5.2, GPT-5.5, and Opus 4.6, respectively. Across all nine configurations, rates range from 16.1% to 62.4%. Our Consequence-Isolated Deterministic Back- end (CIDB) standardizes tool feedback and suppresses downstream effects while retaining the evaluated agent’s model, harness runtime, and tool-dispatch path. The resulting benchmark, CLIGALE, measures prohibited requests and generated content under a shared execution boundary. A traceable Entry × Technique × Impact specification fixes the behavioral workload; separate seed-transfer and judge- aware adaptive tiers extend the analysis to stronger attack preparation. The frozen comparisons show why agent safety reports should identify the complete model–harness configuration and why changing either part warrants renewed evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.