Releasing Predicted Probabilities: Sharp Limits on Excess Disclosure
Abstract
Releasing predicted class probabilities can reveal sensitive information beyond what optimal prediction requires. We show that perfectly calibrated forecasts can approach the optimal expected log loss while revealing a sensitive bit independent of the true conditional class probabilities. We measure this excess disclosure as the largest reduction in optimal expected loss from observing the release instead of the true probabilities, over decisions about other attributes with losses in . We then determine the prediction cost of limiting this disclosure through randomization. For binary forecasts with log-loss regret at most , the smallest worst-case bound on both prediction regret after release and excess disclosure is as . Uniform noise after an arcsine square-root transformation attains this bound, and we prove a matching lower bound for all randomized releases. For any fixed number of classes, releasing counts of labels sampled independently from each forecast achieves the same cube-root rate. In experiments on retinal images and chest radiographs, this sampling method makes images from the same patient harder to link, at the cost of higher average prediction loss when each task uses only its own released probabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.