acceptodds
Under review as a conference paper at ICLR 2027

Hidden in the Features: Semantic Backdoor Watermarking for Large Vision-Language Models

Abstract

Large vision-language models (LVLMs) have become highly valuable intellectual property, making them prime targets for unauthorized copying and downstream reuse. Backdoor watermarking offers a promising route to copyright protection; however, existing LVLM watermarks typically encode ownership signals as specific output patterns or token-level statistics. Such signals produce distribution-shifted responses that are easy to detect, and their effectiveness can be weakened by downstream adaptation. In this work, we present SeMark, a stealthy semantic backdoor watermarking method for visual encoders in LVLMs. Specifically, SeMark employs a bounded universal perturbation as an imperceptible trigger and trains the watermarked encoder to steer triggered representations toward a target semantic neighborhood in the visual feature space. The watermark is embedded into the visual encoder via a bilevel optimization that rigorously balances feature alignment, clean representation preservation, and pairwise similarity. This ensures that triggered features approach the target semantic region, while clean representations remain close to those of the original encoder. Ownership is then verified via black-box queries by assessing whether private triggered samples elicit target-consistent responses from a suspect model. Extensive experiments demonstrate the effectiveness of SeMark, as well as its robustness against downstream fine-tuning while preserving clean-task performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.