Auditing Context-Dependent Safety After Persona-Conditioned Fine-Tuning
Abstract
Safety evaluations commonly attach one score to a fine-tuned checkpoint, although the same weights may be deployed under different system contexts. We audit this assumption in one open 8B model and LoRA pipeline. Models are left untrained or adapted for three rounds with a compliant persona, a principled persona, or generic instruction tuning, then each frozen checkpoint is evaluated through four lenses on the same 313 forbidden prompts. Both persona-conditioned arms retain high refusal under a standard lens and fall under a permissive lens, including a held-out paraphrase; the generic control does not move in the same direction, and a co-primary harmful-usefulness evaluator changes correspondingly. The measurement audit materially changes how this observation can be read: an ASCII-only apostrophe pattern creates treatment-correlated refusal error, and a raised generation cap shows that excluding apparently truncated responses can delete learned behaviour rather than remove artefact. Preregistered validation found the corrected instrument missed no human-labelled refusal and over-fired twice in 283 records, and a second validation over legacy-cohort controls confirms that all 39 correction-flipped records in its four audited seed-42 cells are genuine refusals. Corrected regex–human agreement is ; one author labelled throughout, with pass-1/pass-2 refusal test–retest , so we make no inter-rater claim. We therefore treat safety as a property measured on a checkpoint–context pair, not a checkpoint alone. This is a bounded case study, not a causal estimate of persona semantics: target breadth, response style, prompt authority, update method, and model family remain open.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.