acceptodds
Under review as a conference paper at ICLR 2027

Safety-Constrained Test-Time Training: Post-Update Guarding for Aligned Language Models

Abstract

Test-time training (TTT) adapts a language model during inference, but an attacker can use the same update loop to erode safety alignment. We propose Safety-Constrained Test-Time Training (SC-TTT), which probes each candidate update against a fixed behavioral reference. It retains the candidate only when the harmful-side compliance change does not exceed the benign-side change by more than a threshold. Otherwise, SC-TTT restores the preceding parameter state. Across Gemma-3-4B, Llama-3.2-3B, and Qwen3.5-4B, SC-TTT lowers mean attack success from 19.8-50.3% under vanilla TTT to 0.1-5.8% over five random seeds. It retains 0.83-1.09 of the vanilla adaptation gain on BBH-H3 and matches vanilla accuracy on ARC-AGI. A learning-rate analysis shows that the criterion is reliable only within a calibrated working regime. A full-knowledge three-step pressure test still suppresses 47-70% of vanilla attack success despite optimizing the gate directly. These results establish strong attack suppression, high utility retention, and an explicit operating regime for guarded test-time adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.