Label-Prior Leakage in Federated Evaluation and an Inference Management Fix
Abstract
Federated learning enables collaborative medical AI without sharing patient data, yet post-deployment inference and evaluation remain largely unmanaged. We identify a previously overlooked label-prior leakage in personalized federated evaluation: client-matched testing of private heads implicitly reveals client identity and its label distribution, inflating apparent performance even for a pure prior-only predictor. To address both training heterogeneity and reliable deployment, we introduce FedGIM, a proximal-regularized method with private heads and class-balanced local objectives, together with Federated Inference Management (FedIM). FedIM is a model-agnostic post-training layer that coordinates four stages: Adapt (entropy-minimizing test-time adaptation), Monitor (MMD and Kolmogorov-Smirnov drift detection), Act (confidence-gated computational escalation, optional differential-privacy logit release, and model fusion), and Refer (selective classification with explicit coverage control and human referral). Evaluated on non-IID endoscopic screening under a leakage-aware pooled protocol with label-free feature-space routing, FedGIM recovers near-centralized accuracy while FedIM provides controllable coverage-risk trade-offs, drift response, latency-accuracy balance, and privacy-utility curves. Our results demonstrate that evaluation protocol choice can dominate reported gains and that managed inference is essential for clinically trustworthy federated systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.