acceptodds
Under review as a conference paper at ICLR 2027

Robustly Contestable Ownership Proof from Model Memorization

Abstract

Open-weight models can be adapted and redistributed: Alice trains and releases a model, Bob then fine-tunes it but claims he trained it himself. How can Alice provide a publicly verifiable proof that Bob’s model is derived from hers? We study this goal as a contestable ownership proof: Alice wants to prove that she trained her model and wants to contest any declared lineage that omits her release. Despite progress in model fingerprinting, which embeds detectable signals in models, existing methods do not enable Alice to publicly prove that Bob’s model was derived from hers. In this paper, we present the first such construction: a generic framework that uses cryptography to bind a model fingerprint to its developer’s identity and declared lineage, enabling any third party to verify ownership proofs and lineage challenges. We instantiate the framework with palimpsestic memorization, the tendency of models to memorize later training data more strongly. We implement the resulting protocol as a Python library. To test the robustness of this instantiation, we study adversarial fine-tuning that aims to make the original developer’s ownership claim fail verification. We evaluate six models of different sizes from three families: OLMo, Pythia, and CrystalCoder. For each model, we use linear programming to find a fine-tuning direction predicted to reduce the fingerprint signal as much as possible. Along these tested directions, the ownership test stops detecting the signal only after WikiText perplexity rises substantially.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.