LoRo-Mark: Provably Lossless And Robust Agent Watermarking
Abstract
As large language model (LLM) agents are increasingly deployed as commercial services, protecting their proprietary orchestration logic and tool-use policies has become an important concern. We consider a realistic infringement scenario, termed **agent repackaging**, in which an adversary integrates a protected agent into its own application through an API and presents it under its own service identity. The adversary may further modify parts of its execution process to obscure the original source. In such cases, the owner typically has access only to the repackaged service interface, making black-box ownership verification essential. Agent watermarking provides a natural way to embed ownership evidence into the agent’s behavior for later verification. Under this setting, we argue that an effective watermark should satisfy two key requirements: losslessness, requiring watermarking to preserve the agent’s original functionality, and robustness, requiring ownership evidence to remain recoverable after partial modification of the execution process. Existing methods, however, typically embed ownership signals into behavior selection or execution trajectories, intervene in normal agent decisions and provide limited robustness to behavior modification. To address these limitations, we propose **LoRo-Mark**, a provably **lo**ssless and **ro**bust agent water**mark**ing mechanism. For losslessness, LoRo-Mark isolates watermarking into a cryptographically authenticated forensic branch that remains inactive during normal execution and can only be activated by owner-authorized requests. By reducing unauthorized branch activation to standard MAC security, LoRo-Mark provides a formal cryptographic guarantee of performance preservation. For robustness, it redundantly distributes ownership information across forensic behavior sequences, enabling reliable recovery under partial behavior substitution and sequence truncation. Experiments across multiple LLM agents demonstrate zero degradation on normal tasks and reliable ownership verification under sequence modifications.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.