Selecting Red-Teaming Policies under Implementation Shift
Abstract
An adaptive red-teaming policy can rank first on one implementation of a hidden behavior and fall behind on another, even when the trigger, access, and query bud- get are held fixed. We study this selection problem through a reference intervention: a twin system keeps the target’s on-trigger responses and replaces its off-trigger responses with those of a fixed reference. The twin inherits the reference’s first-hit rate, which isolates the role of implementation-dependent feedback. A change-of- measure argument stopped at the triggering query bounds the resulting localization shift using only the off-trigger responses that precede it, and a justified upper bound on this information budget yields target-rate intervals and conditional ranking certificates. Localization alone does not identify discovery order: we construct opposite discovery rankings with zero localization shift and give guarantees under additional lower bounds on first-hit recognition. On Qwen3-8B and Gemma-3-12B- IT, choosing between two policies on the reference rather than on a prompt-based implementation improves recognized discovery on a held-out trained implementa- tion by +11.82 and +6.54 percentage points, respectively (one adapter per model). Controlled targets reverse this advantage when the held-out implementation retains the diagnostic feedback of the implementation used for selection. A finite-channel experiment further shows that statistically valid calibration can cost more queries than it saves in target validation. These results identify when reference evaluation supports policy selection and which additional assumptions are needed to certify it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.