acceptodds
Under review as a conference paper at ICLR 2027

LLM Agents Can Easily Tamper Their Own Traces

Abstract

Asynchronous monitoring, incident investigations, and compliance audits rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own traces. We show that almost all local LLM agents like Claude Code and Codex fail to enforce this boundary. All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails. We also validate that external attackers can exploit this gap to induce trace deletion. Finally, we show that trace tampering behavior emerges naturally in frontier models, when agents try to improve their rewards. We advise practitioners to ensure trace logging happens through an independent interception mechanism outside of the agent's control, preserving trace integrity even in cases of full host compromise. Overall, our findings identify a concrete failure of trace integrity in agent infrastructure which can be used to conceal misaligned behaviors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.