Inferring Large Language Model Provenance via Token Ranking
Abstract
Model provenance detection determines whether a model was derived from a claimed parent through fine-tuning or other adaptations. Many methods infer provenance relationships by measuring the similarity between two models from their generated responses. However, these approaches rely solely on sampled tokens, leaving the rich information associated with alternative candidate tokens unexplored. In this paper, we propose **RankMPT**, a novel method that leverages the rank information of candidate tokens to detect model provenance. Specifically, RankMPT measures the average rank correlation between the tail candidate tokens of two models across the prompt tokens, which can be approximated via repeated high-temperature sampling in black-box settings. We then perform a hypothesis test to decide whether the score is significantly higher than the correlation level expected in the absence of a provenance relationship. Our method captures discriminative signals from rarely selected tail tokens, making it less sensitive to confounding similarities among models arising from shared domains or tasks. Extensive experiments demonstrate the superiority of RankMPT, improving the average accuracy by up to 23% in the black-box setting of LeafBench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.