acceptodds
Under review as a conference paper at ICLR 2027

From Textual Outputs to Probabilistic Signatures: Robust Black-Box Model Ownership Verification

Abstract

The substantial computational resources required to train large language models (LLMs) make them valuable intellectual property, raising concerns about model theft and unauthorized deployment. Fingerprinting provides a practical approach to ownership verification by determining whether a suspicious model is derived from a protected source model. However, existing black-box methods often fail to provide reliable verification, as textual signals are unstable under stochastic sampling. This issue is pronounced for reasoning models, where long-form generation amplifies the variability introduced by sampling. Although sampled texts vary, the underlying probability patterns remain relatively stable. Motivated by this insight, this work proposes Token-Level Probability Fingerprinting (), a robust framework that leverages probability distributions as more stable verification signals than textual outputs. Since derived models tend to preserve probability patterns similar to their sources, extracts discriminative features by evaluating the token-level probabilities assigned by the source model to responses generated by a suspicious model. Extensive experiments on both reasoning models and standard LLMs demonstrate that achieves effective verification performance approaching white-box methods. Code is available in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.