SynChain: Inducing Computer-Use Agent Systems to Construct Their Own Attack Chains
Abstract
Computer-use agents (CUAs) have turned large language models into persistent execution systems that generate, store, and reuse their own artifacts, such as skills and memory entries. Existing security defenses, however, largely treat attacks as externally triggered and temporally bounded, overlooking how compromise can propagate internally through an agent's own persistent state. We show that malicious influence can be covertly embedded into the structural redundancies of self-synthesized artifacts, where it survives internal state updates and evades standard vetting. We formalize this threat as , a self-synthesized attack paradigm that uses persistence-aware directed supervised fine-tuning to induce agents to produce poisoned yet benign-looking artifacts. The resulting dormant payloads later reactivate as trusted context in future workflows, without any new malicious external input. To evaluate this propagation at scale, we build CUAChain, a chain-structured benchmark covering five categories of real-world agent workflows, multi-step task chains of up to five subtasks, and three attack objectives, and conduct over 6,000 end-to-end agent executions. Across OpenClaw, Codex, and Claude Code under four defense settings, SynChain achieves over 93% attack success in single-hop propagation and consistently outperforms adapted baselines, showing that securing CUAs requires provenance-aware reasoning over cross-task execution trajectories.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.