Hide in Plain Sight: Towards Scalable and Text-only LLM Copyright Auditing via Watermarking Token-wise Attribution
Abstract
Safeguarding the intellectual property of Large Language Models (LLMs) has become a critical priority. While black-box watermarking offers a proactive solution, existing paradigms face significant limitations: output-space methods inevitably compromise model utility and functional integrity, whereas current alternative-space techniques rely on fine-grained logits and computationally expensive global regression. To address these challenges, we propose CAST, a scalable and text-only auditing framework via watermarking token-wise attribution. Specifically, CAST decouples the embedding process into independent, token-wise attribution and anchors the watermark signal to the importance scores of input tokens on a vocabulary subset . This framework fundamentally resolves the memory bottleneck by achieving a length-independent space complexity and enables reliable text-only verification, where signals are recovered via frequency-based statistical estimation. Extensive experiments on various modern LLMs demonstrate that CAST achieves near 100% watermark success rates in text-only settings while preserving the original model's utility and scalability, significantly outperforming state-of-the-art LLM watermarking methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.