Output Probabilities Under Homomorphic Encryption: Depth Chosen for Accuracy Is Not Enough
Abstract
Homomorphic encryption (HE) lets a server run a model on data it never sees, and HE inference is judged by task accuracy, approximation error, and multiplicative depth. Many deployments, however, consume the output probabilities themselves, through confidence thresholds, abstention rules and cost-weighted decisions, and none of those three quantities certifies them. Replacing the output softmax by a polynomial, the usual HE-compatible route, leaves the predicted label unchanged in every seed of ten backbone and dataset settings while inflating expected calibration error on the same logits by 1.1x to 30x, and by roughly 10x to 20x over the full vocabularies of two language models. Under the repeated-squaring schedule used for high-precision homomorphic exponentials, accuracy reaches its ceiling one to four multiplicative levels before calibration comes within 0.005 of the exact softmax: the multiplicative depth sufficient to preserve predictions can be systematically insufficient to preserve output probabilities. We trace the effect to a non-affine residual in pairwise log-odds, which yields a label-free diagnostic, computable from the public polynomial and the operating range before deployment, that chooses depth where accuracy cannot. The repair depends on the interface. Where logits may leave the ciphertext, defer the softmax to the client; where an existing pipeline already returns polynomial numerators, invert them in plaintext as a compatibility layer, using a numerator monotone by construction; where the distribution itself must stay encrypted, fold the calibration temperature into the model and let the diagnostic set the depth, which in exact-arithmetic simulation lands within 0.004 ECE of the client-side result on 96 of 99 seeds. Real CKKS runs of the head and numerator reproduce the simulated effect.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.