acceptodds
Under review as a conference paper at ICLR 2027

From Bits to Beliefs: Recovering Owner Identity for Black-Box Verification of Large Language Models

Abstract

Open-weight large language models (LLMs) can be copied, modified, and redeployed behind black-box APIs, making post-release ownership verification difficult. Existing black-box fingerprinting methods largely treat verification as a detection problem, asking whether registered query-response behaviors remain observable in a suspect model. Such evidence can degrade under fine-tuning, pruning, quantization, model merging, and serving-time prompt changes. We propose SimPrint, a recoverable semantic fingerprinting framework that instead formulates ownership verification as identity recovery from noisy black-box behavior. SimPrint encodes a private owner identity into an error-correcting codeword and distributes its bits across natural binary question-answering probes. It selectively implants only probes that must deviate from the base model through a low-interference batch update, and later parses suspect-model responses into bits or erasures to recover the registered identity. Because verification requires only input-output queries, SimPrint remains applicable when model weights and activations are inaccessible. Experiments on three open-weight LLMs show reliable identity recovery in both clean and modified settings, robustness to fine-tuning, pruning, quantization, model merging, and serving-time perturbations, and comparable downstream utility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.