acceptodds
Under review as a conference paper at ICLR 2027

Memory Control Shapes Adversarial Sensitivity and Specialization in Recurrent Language Models

Abstract

We show that recurrent memory control shapes not only how much adversarial damage is produced, but also which perturbations are effective. Using Gated DeltaNet-2 (GDN2), which explicitly decouples memory erase from write, together with corresponding controls over new information input and old state decay in Mamba-1 and Hymba, we study these effects under adaptive white box attacks. First, write control has a substantially larger effect on adversarial damage than erase across all three architectures. This suggests that inputs which naturally induce stronger write activity may allow adversarial perturbations to have greater downstream influence. In GDN2, this ordering is opposite to the clean performance ablation trend, where erase plays the larger role. Second, changing memory dynamics can weaken a fixed attack without reducing vulnerability to an adapted attack. In our GDN2 post attack erase experiment, about 95% of the apparent recovery comes from the original attack transferring poorly to the modified memory condition, rather than from lower damage under an adapted attack. Finally, we show how this loss of transfer develops during optimization. As optimization proceeds, attacks become stronger on their source memory condition but transfer less effectively to other conditions. Along the same trajectories, gradients under different memory conditions become less aligned and the corresponding optimized perturbations become less similar. We refer to this process as memory condition specialization. Overall, recurrent memory dynamics shape both adversarial sensitivity and attack specificity, highlighting the need to adapt attacks to the target memory condition when evaluating robustness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.