Functional Network Fingerprint: A Unified Method for Cross-Architecture and Cross-Structure LLM Lineage Verification
Abstract
Training large language models (LLMs) is extremely expensive, yet released weights can be copied and re-released by third parties. We study a subtle disguise, which we call weight repackaging: an attacker copies a victim model's weights and then alters its structure—adding, removing, or interpolating layers, or expanding the hidden or MLP width—before continuing training. The resulting model differs from the victim in depth, width, and tokenizer while inheriting most of its parameters, making fingerprints that rely on direct weight or representation correspondence difficult to apply. We propose the Functional Network Fingerprint (FNF), a white-box method that requires no fingerprint-specific training and verifies whether a suspect model is derived from a victim model. Inspired by brain-fingerprinting research, FNF applies canonical independent component analysis (CanICA) to internal activations of Transformer blocks to identify co-activating neuron groups, and compares their activity trajectories between two models under identical input stimuli. Because FNF compares functional activity rather than weights or static representations, it does not require direct correspondence in depth, width, or tokenizer, and can operate across architectures and under common weight transformations. On a labeled set of model pairs spanning three provenance categories, FNF separates derived from independent models with an AUC of and identifies every derived model at zero false positives. On the pairs to which two non-invasive baselines also apply, the representation-based method REEF reaches an AUC of and the recent attention-difference method AttnDiff , and at zero false positives they recover only and of the related pairs. FNF thus provides a simple and interpretable tool for LLM provenance verification.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.