Benign Safety Is Not a Robustness Certificate: Evidence of a Structural Evaluation Gap Across Attack Conditions
Abstract
We study whether safety evaluation under benign conditions can substitute for direct adversarial robustness evaluation. We operationalize benign evaluation using a specified, equally weighted composite of helpfulness, harmlessness, and honesty across six selected instruction-tuned checkpoints. Exact-token assistant-prefill trajectories reveal the first level of measurement non-interchangeability: stronger scores on this composite do not consistently imply stronger prefix robustness. Among the original four checkpoints, robustness ordering under the prefix condition does not fully carry over to a structurally different transferred adversarial condition, establishing a second finite failure of substitution. Repeated interaction exposes additional frozen-evaluator failures missed by first-response assessment. Within the prefix protocol, model ordering remains stable across the tested grids even though robustness magnitudes depend materially on grid resolution. Human calibration with independent primary author annotation and third-author adjudication provides moderate support for the harm evaluator, with weaker agreement on adaptive examples. We characterize these findings as an empirical structural evaluation gap, not a defect in alignment-training mechanisms. They motivate distinct, condition-explicit benign and adversarial measurements, without implying that every benign evaluation or adversarial condition shares the observed relationships.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.