SafeLoop: Preserving Refusal of Harmful Prompts in Looped Language Models
Abstract
Refusing to fulfill harmful requests is a crucial safety behavior of language models. Looped language models offer a parameter-efficient approach to reasoning by reusing the same Transformer weights for additional computation before predicting each token. However, whether refusal remains reliable as recurrent depth increases is not well understood. Evaluating this reliability through attack success alone can be misleading: deeper models can stop producing harmful answers because their generation has degenerated. We identify a distinct failure in refusal representations. A behaviorally validated refusal direction can recover harmful–benign separability after erasure in one loop, yet its projection can drift far from its reference value during unperturbed inference. The ability to reconstruct this representation therefore does not ensure that additional loops preserve it. This finding motivates SafeLoop, which stores each input's projection at a reference loop and corrects subsequent deviations without updating model parameters. On frozen Ouro-1.4B, SafeLoop increases the percentage of responses that are non-degenerate refusals (effective refusal) from 9.5% to 41.0% at 16 loops and from 6.5% to 35.5% at 24 loops. It also improves effective refusal in selected recurrent models converted from dense backbones. Benefits vary across models and depths; projection correction does not resolve widespread generation degeneration and can increase benign over-refusal. These results establish cross-loop preservation as a concrete target for refusal control and show why evaluating deeper inference requires separating effective refusal from generation failure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.