Adaptation or Poisoning? Security Risks of Conversational Test-Time Learning in Large Language Models
Abstract
Test-time learning (TTL) enables large language models to adapt to unlabeled inputs during deployment. The model updates a lightweight adapter by minimizing the next-token prediction loss on the input it receives. When these updates persist across conversation turns, every user query serves both as the request answered in the current turn and as adaptation data for later turns. By choosing the queries used for adaptation, an attacker can influence parameter updates and weaken safety alignment through ordinary interactions with a TTL chat service. To demonstrate this risk, we introduce Progressive Test-Time Adaptation Poisoning (P-TAP), a query-only multi-turn attack that jailbreaks the model by poisoning its session adapter through the service’s own updates. A sequence of escalating adversarial queries shapes the adapter turn by turn; the final harmful request is then answered under the accumulated state without itself being used for adaptation. Across four open-weight instruction-tuned models, enabling updates raises attack success rates on average by 14.50 percentage points under greedy decoding and by 24.25 points under best-of-10 sampling, relative to matched sessions receiving the same queries without updates. Answering adversarial queries but replacing them with unrelated benign text for adaptation removes most of this increase. Among existing safeguards, screening incoming queries flags every attack stream before the final request, while a perplexity-based detector of unsafe adaptation flags none of the successful attacks. the optimizer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.