acceptodds
Under review as a conference paper at ICLR 2027

A Released LoRA Adapter Names the Document It Was Trained On

Abstract

LoRA adapters are low-rank factors smaller in size compared to the base models they are fine-tuned on. Over 130,000 LoRA adapters are publicly posted on Hugging Face alone. Anyone who downloads an adapter receives it without knowing the data it was fine-tuned on. Prior studies of these weights have produced a scalar or a verbal description of the training data, plus a generative attack reports ROUGE-2 below 9% on the source text, thereby concluding that extraction is extremely difficult. In this work, we show that a released adapter identifies the exact document it was fine-tuned on. Our method retrieves that document from roughly 570k candidates for all 256 held-out adapters trained for five steps and shows a 0.01% false positive rate across millions of shortlisted negatives. In that same setting, greedy decoding reproduces no document verbatim, so identification does not require the text to be recoverable. On a second base model, we refit the method end-to-end and it again identifies every document, while decoding reproduces at most one of 256. We trace the signal to the gauge-invariant product . A kernel over dW alone with no base model, tokenizer or forward pass, ranks the training document at C-AUC 0.897 among 16 candidates. We also measure the bounds of this method. We see that by eighty steps greedy decoding can recover the document and training one adapter on several documents dilutes which one is named. More broadly, our work shows that publishing an adapter can disclose the document behind it and privacy assessments must extend from a model's outputs to its weights.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.