Reducing Training-Value Leakage in Split Inference via Learned Weight Perturbations
Abstract
Split inference avoids transmitting raw inputs, but the intermediate representations exposed at the split point may still reveal valuable information. In proprietary industrial, scientific, and other high-value data settings, the primary concern is not only whether individual samples can be visually reconstructed, but whether the exposed information can be recovered and reused to train another model. We refer to this risk as training-value leakage. To mitigate it, we propose learned client-side weight perturbations that modify selected pre-split weights while leaving the server model and deployment interface unchanged. The perturbations are optimized with reconstruction-aware feedback and selected under an explicit validation-accuracy constraint, so that reduced leakage is not obtained by simply degrading the original prediction task. We evaluate the method on CIFAR-10, CIFAR-100, and NEU-CLS using post-selection SplitNN-decoder and UnSplit reconstruction benchmarks, and measure reconstruction MSE and reconstructed-dataset trainability across three evaluator architectures. Learned perturbations reduce split-inference accuracy by only 2.3 percentage points on average, while reconstructed-dataset trainability shows an average positive drop of 23.2 points and a maximum drop of 65.7 points. These results suggest that split representations can retain intended inference utility while becoming substantially less reusable as downstream training data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.