From Textual Outputs to Probabilistic Signatures: Robust Black-Box Model Ownership Verification
Abstract
The substantial computational resources required to train large language models (LLMs) make them valuable intellectual property, raising concerns about model theft and unauthorized deployment. Fingerprinting provides a practical approach to ownership verification by determining whether a suspicious model is derived from a protected source model. However, existing black-box methods often fail to provide reliable verification, as textual signals are unstable under stochastic sampling. This issue is pronounced for reasoning models, where long-form generation amplifies the variability introduced by sampling. Although sampled texts vary, the underlying probability patterns remain relatively stable. Motivated by this insight, this work proposes Token-Level Probability Fingerprinting (), a robust framework that leverages probability distributions as more stable verification signals than textual outputs. Since derived models tend to preserve probability patterns similar to their sources, extracts discriminative features by evaluating the token-level probabilities assigned by the source model to responses generated by a suspicious model. Extensive experiments on both reasoning models and standard LLMs demonstrate that achieves effective verification performance approaching white-box methods. Code is available in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.