SAME: Semantic Watermark Embedding for Open-Source Large Pretrained Models
Abstract
Open-source large pretrained models (LPMs) have become increasingly valuable assets for model development and downstream applications, but their accessibility also facilitates unauthorised copying and modification. To protect their intellectual property (IP), existing approaches typically embed tailored trigger–response patterns as watermarks for ownership verification. However, such trigger designs need to be tailored to task-specific data that constrain their transferability across downstream tasks, and are not durable under subsequent benign fine-tuning. To address these limitations, we propose SemAntic waterMark Embedding (SAME), which introduces a semantically meaningful private vocabulary that enables flexible composition of adaptive and persistent watermarks. Specifically, SAME introduces a fixed private vocabulary that assigns each native token a private embedding counterpart, while aligning the semantic expressiveness of the two vocabularies within the model’s representation space. This alignment allows private tokens to be flexibly composed into semantic triggers that elicit meaningful responses, enabling adaptive watermarks beyond predefined trigger–response pairs. By grounding the private vocabulary in the model’s native semantics, SAME couples the watermark with its evolving semantic capabilities, allowing this connection to be preserved throughout downstream fine-tuning. Extensive experiments demonstrate that SAME provides reliable and robust watermarks across LPMs of varying scales and architectures, enabling effective ownership verification for IP protection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.