acceptodds
Under review as a conference paper at ICLR 2027

ContamMorph: Do Transformed Views Add Membership Evidence After Fine-Tuning?

Abstract

Membership inference asks whether a passage was used to train a model. Auditors can average scores across reformatted or rewritten versions, but a detection gain may come from one transformation alone. We compare the original text, a fixed transformation, and their average to separate these effects. We randomly include or withhold short English passages during fine-tuning, using two corpora and models with 0.5 to 1.5 billion parameters. Our main case study is error-zone membership inference (EZ-MIA), which divides accumulated prediction improvements by accumulated losses. Line breaks improve this detector before averaging, largely through the released code's handling of a zero denominator. When some incorrectly predicted text pieces improve and none worsen, the code assigns zero, whereas the published rule ranks the passage first. One shared positive denominator constant restores strong detection in four model and corpus settings and in fresh training runs. At an intermediate training strength, the implementation choice moves EZ-MIA from the top two of six detectors to near the bottom. After correction, averaging rarely beats an original-text detector selected on other runs. We observe this pattern in twelve-run replications, when included passages form 5% of a larger training set, and with five separately generated paraphrases. Those paraphrases distinguish members from controls less well than the exact wording used in training, and their average performs worse. Auditors should compare every aggregate with both the original text and a prespecified transformed version.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.