StageBreak: Diagnosing Stage-Selective Safety Failures in Multimodal Language Models
Abstract
Jailbreak evaluation usually reduces multimodal safety to a single endpoint: whether an unsafe response is produced. That endpoint conflates three different cases. The model may fail to recognize risk, lose an early refusal preference, or cross the safety boundary later in decoding. We introduce StageBreak a diagnostic framework that assigns the earliest valid operational failure among risk recognition, refusal preference, and safe generation while preserving the required upstream criteria. A stage attribution is accepted only for clean-eligible examples whose risky task semantics remain invariant and whose harmful endpoint, prerequisite-preservation, and audit gates agree on the same sample. Low-rank repair, reverse transplantation, cross-stage controls, and cross-attack transfer then measure the interventional relevance of the attributed representation. The resulting attribution distinguishes semantic corruption from recognition, refusal-preference, and generation-time failures. We evaluate whether these attributed failures persist across tasks, external benchmarks, and held-out models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.