acceptodds
Under review as a conference paper at ICLR 2027

MemMorph: Persistent Tool Hijacking in LLM Agents via Long-Term Memory Poisoning

Abstract

Tool selection is security-critical in LLM agents: selecting an inappropriate tool can expose sensitive information, trigger unauthorized actions, or derail intended workflows. Existing tool-hijacking attacks typically manipulate tool descriptions or metadata, embedding the attack payload directly in the tool interface. We propose MemMorph, a memory-poisoning attack for persistent tool hijacking that leaves the tool interface, model parameters, and system prompt unchanged. MemMorph injects a few crafted memories in factual, episodic, and procedural forms that bias tool choices through retrieved context. Because long-term memory modules may summarize or rewrite records before persistence, MemMorph optimizes poisoned memories for future retrieval while constraining them to retain their effect on tool selection when such rewriting occurs. Across nine target scenarios spanning three benchmarks and ten downstream LLMs, three poisoned memories achieve 59.2-78.1% target-tool selection, outperforming the strongest memory-poisoning baseline by 13.4 points on average. MemMorph remains effective under indirect injection, transfers across alternative memory processors and modules, and continues to induce target-tool selections as the memory store grows 6.7 after a one-time injection. Our findings show that long-term memory can serve as a persistent channel for tool hijacking in tool-using LLM agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.