acceptodds
Under review as a conference paper at ICLR 2027

Benign Safety Is Not a Robustness Certificate: Evidence of a Structural Evaluation Gap Across Attack Conditions

Abstract

We study whether safety evaluation under benign conditions can substitute for direct adversarial robustness evaluation. We operationalize benign evaluation using a specified, equally weighted composite of helpfulness, harmlessness, and honesty across six selected instruction-tuned checkpoints. Exact-token assistant-prefill trajectories reveal the first level of measurement non-interchangeability: stronger scores on this composite do not consistently imply stronger prefix robustness. Among the original four checkpoints, robustness ordering under the prefix condition does not fully carry over to a structurally different transferred adversarial condition, establishing a second finite failure of substitution. Repeated interaction exposes additional frozen-evaluator failures missed by first-response assessment. Within the prefix protocol, model ordering remains stable across the tested grids even though robustness magnitudes depend materially on grid resolution. Human calibration with independent primary author annotation and third-author adjudication provides moderate support for the harm evaluator, with weaker agreement on adaptive examples. We characterize these findings as an empirical structural evaluation gap, not a defect in alignment-training mechanisms. They motivate distinct, condition-explicit benign and adversarial measurements, without implying that every benign evaluation or adversarial condition shares the observed relationships.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.