Who Made This LLM? Black-Box Family Provenance for Large Language Models
Abstract
Frontier large language model (LLM) families such as Llama, Qwen and Mistral are valuable assets, and the ease of producing derivative models through fine-tuning, distillation or merging exposes them to theft, unauthorized reuse and false ownership claims. Identifying the family a deployed LLM descends from is therefore central to ownership auditing, but existing solutions fall short under realistic conditions. Watermarking requires proactive deployment and cannot be retrofitted to already-released models, white-box fingerprinting requires parameter access that API-served models do not expose, and extracting reliable family-level signals from black-box responses remains challenging under model modifications. We propose FamilyTrace, a black-box fingerprint for family-level LLM provenance that supports both closed-set attribution among known families and open-set rejection of unknown families. Our design starts from two observations: much of the knowledge a model learns in pre-training is encoded in its feed-forward layers, which hold most of its parameters, and typical post-training changes these weights only slightly. We therefore expect the models of a family to share similar word associations, and FamilyTrace captures them by asking a model to rank candidate words by their relation to an ambiguous word, first with little context and then with progressively more context. We evaluate FamilyTrace on two benchmarks covering six LLM families, a non-adversarial benchmark of 55 LLMs and an anti-provenance benchmark of 50 LLMs. FamilyTrace achieves 96.3% closed-set and 87.9% open-set accuracy on the non-adversarial benchmark, and 85.4% closed-set and 79.5% open-set accuracy on the anti-provenance benchmark, outperforming seven black-box fingerprinting baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.