Technique Awareness as an Alignment Risk: Model Familiarity of Inoculation Prompting Weakens its Effect
Abstract
Natural emergent misalignment (EM) shows that models which learn to reward hack during capability RL can generalize into broadly misaligned behavior. Inoculation prompting (IP) — e.g., telling models hacking is okay — largely neutralizes this, and is now part of production RL recipes for frontier models. However, IP rests on a lie (developers don't actually want hacks), which presumes the model takes it at face value. **With future models knowing more about IP, we ask whether such *technique awareness* weakens the inoculation.** **The answer is yes.** A Llama-3.3-70B-Instruct given lightweight IP-teaching synthetic document finetuning before RL ends up emitting 2.4x the natural EM readings of its IP-unaware counterpart. Surprisingly, IP-aware models are also far more misaligned when IP is absent during RL, making such awareness tricky knowledge for a model to hold. We further find IP-aware models can exhibit “no-hack EM” without ever being rewarded for hacking. Together, these make aligning IP-aware models a significant challenge: applying IP helps less, withholding it hurts more, and even a hack-free environment no longer implies safety. We employ concept ablations and large-scale reasoning trace analysis to identify what contributes to these observations. We also opensource the first recipe, data, and checkpoints that reliably elicit natural EM and demonstrate IP working — a setting so far confined to internal frontier lab work and repeatedly reported as hard to reproduce on an opensource stack — which we envision as a useful testbed for studying misalignment under more realistic threat models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.