The Calibration Set Is in the Checkpoint: How Release Geometry Governs Membership Leakage in Post-Training Quantization
Abstract
Quantizing a model often uses a small calibration set after training is complete. This set can remain encoded in the released checkpoint. In SmoothQuant-style methods, comparing the released checkpoint with its public base model recovers activation statistics computed from the calibration data. A maximum can copy a calibration record's activation exactly into the checkpoint, while averaging removes this exact constraint but can keep other membership signal. We show that leakage grows on a common scale, the number of released channels per calibration record (), while the released statistic sets its shape. On OPT-125M, the reference SmoothQuant implementation and NVIDIA Model Optimizer's stock exporter each identify about 68% of calibration records at measured false-positive rates below . Changing only the statistic in a Qwen2.5-7B exporter changes detection at low false-positive rates by an order of magnitude at essentially unchanged perplexity, and the same effect appears in isotonic probability calibration. Post-training calibration can therefore itself be a privacy-relevant release mechanism.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.