Fooled by Fine-Tuning: Why Memorization Audits of Code-Completion Models Need Controls
Abstract
Memorization audits of code models commonly inject canaries into training data and measure how strongly the resulting model prefers them, often without evaluating comparable non-members. We show that this practice can produce convincing but misleading conclusions. Through a controlled audit of fine-tuned code models with more than canaries, multiple attacker families, checkpoints, and controls, we identify three recurring failure modes. First, large reference-model exposure scores do not necessarily distinguish injected canaries from never-injected controls, even with a matched-corpus reference. Second, a completion-based likelihood-ratio membership test can respond similarly to members and non-members, effectively detecting fine-tuning rather than membership. Third, extraction guarantees can be substantially overstated when confidence bounds are computed for a different estimand than the security claim being made. Positive-control experiments confirm that these findings are not simply failures of the auditing pipeline. We further revisit the information-theoretic certificate motivating our audit. We correct a vacuous additive bound formulation, but show that the pointwise likelihood ratios computed by reference-ratio audits do not directly estimate the divergence required by such a certificate; accordingly, we make no certification claim. Our results motivate a stricter methodology for memorization auditing based on non-member controls, matched references, per-instrument positive controls, and estimand-aware reporting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.