acceptodds
Under review as a conference paper at ICLR 2027

The Privacy Side Effect of Cheap Inference

Abstract

Reducing the numerical precision of model weights is common in deployment because it lowers computational cost. This reduction also selectively lowers a model’s ability to reproduce the training data word-for-word, by several bits, before it substantially affects general capability. This difference is termed the selectivity gap. The gap is measured across a 29-fold range in parameter count and is replicated in a second model family, an error-correcting production quantizer, and text-to-image diffusion, where the gap is more pronounced. Although the gap narrows as models grow, it stays above 24.2 % points of retention across all sizes tested. The content destroyed is not predicted by data duplication, although a magnitude filter, such as quantization, should spare duplicated content. Instead, the sharpness of the loss surface around a memorized sequence predicts its sensitivity to numerical rounding. This extends prior observations linking memorization to sharp loss regions by showing that the same sharpness also predicts which memorized content is lost under quantization. In diffusion models, the selectivity gap persists, but weight-space sharpness is no longer predictive. Instead, the gap is associated with the stability of image generation to perturbations in the starting noise. Finally, decomposing a realistic deployment recipe shows that the 8-bit activation quantizer, rather than the 4-bit weight quantizer, is primarily responsible for the loss of training-image reproduction. Numerical precision is therefore both an efficiency-related deployment choice and a privacy-relevant configuration that can substantially affect a model’s ability to reproduce its training data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.