acceptodds
Under review as a conference paper at ICLR 2027

Scalable Paraphrase-Robust Language Model Fingerprints via Secret Semantic Targets

Abstract

Model fingerprinting aims to let model owners verify, using only black-box API access, whether a hosted service is serving their model or a close derivative. Practical Large Language Model (LLM) fingerprints must satisfy two requirements that prior work has mostly studied separately: they should support many independent fingerprint units, and they should remain detectable under semantic-preserving rewriting such as input and output paraphrasing. Existing methods usually achieve one or the other: scalable memorization-based fingerprints rely on lexical evidence, while statistical watermark-style paraphrase-resistant methods do not provide independently assignable fingerprint units. We introduce scalable semantic fingerprinting via semantic-target fingerprints, where each fingerprint unit is a natural concept paired with a secret target region in embedding space. We instantiate this idea with Semantic Perinucleus Fingerprints, which select plausible low-prior answer modes, and Semantic Cluster Fingerprints, which discover stable completion clusters and train toward selected secret targets. To align training with verification, we use reinforcement-learning to reward the same semantic evidence used at verification time. In our strongest setting, GRPO-trained SCF supports \(K=1024\) simultaneous fingerprint units at a fixed verifier budget of \(N=mK=1024\) (\(m=1\) query per unit), retains \(97.0%\) of Qwen2.5-7B-Instruct measured utility, and achieves a near-perfect normalized verification score. Verification for both SCF and SPF remains intact under input, output, and joint paraphrasing attacks. These results place LLM fingerprinting in the previously missing regime of scalable, paraphrase-robust black-box verification.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.