acceptodds
Under review as a conference paper at ICLR 2027

SinkMark: Robust Model Fingerprinting via Attention-Sink-Induced Hallucinations

Abstract

Large Vision-Language Models (LVLMs) are increasingly released and adapted for downstream applications, making it challenging to verify whether a suspicious black-box model is derived from a copyrighted checkpoint. In this work, we propose SinkMark, a robust model fingerprinting framework that turns model hallucinations from an undesirable failure mode into a behavioral ownership signal. Our key idea is to construct private image-text probes that induce a predefined target hallucination in the source model and its derivatives, while remaining inactive on unrelated models. To achieve this, SinkMark dynamically identifies a shared attention sink across multiple decoder layers, promotes cross-layer agreement and attention concentration at the sink, and binds its hidden representation to the target hallucination semantics. We jointly optimize a bounded image perturbation and a natural-language suffix using PGD and GCG, respectively. To improve persistence under downstream model modifications, we further introduce robust min–max optimization with temporary adversarial parameter updates and predecessor-guided target induction. Importantly, SinkMark does not permanently modify the released model and requires only black-box access during ownership verification. Experiments demonstrate that SinkMark provides reliable and model-specific target hallucinations, while remaining robust to fine-tuning, pruning, and quantization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.