acceptodds
Under review as a conference paper at ICLR 2027

SteerPrint: Steering-Vector-Guided Fingerprints for Black-Box LVLM Ownership Verification

Abstract

Growing concerns over the unauthorized reuse of large vision-language models (LVLMs) highlight the need for black-box fingerprinting techniques that can trace model origins and verify ownership. Existing black-box fingerprints are built mainly around output-level trigger-response associations, with little explicit grounding in internal features inherited from the source model, leaving their persistence under model derivation without a clear internal basis. To overcome this limitation, we introduce and formalize provenance anchors: internal features that persist in derived models while distinguishing them from non-derived models, and further characterize the conditions under which they can be translated into black-box fingerprint signals. Building on this framework, we propose SteerPrint, which uses behavioral directions in representation space as provenance anchors and distills their behavioral effects into visual fingerprints for black-box ownership verification. We evaluate SteerPrint across recent LVLMs, downstream fine-tuning, and model transformations including quantization, pruning, and model merging. The results show that SteerPrint achieves stronger overall verification performance and transferability across derived models than existing black-box fingerprinting methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.