acceptodds
Under review as a conference paper at ICLR 2027

Privacy-Preserving Clinical Speech Embeddings under Multi-Recording Identity-Linkage Attacks

Abstract

Privacy in learned representations can fail under repeated observation. An embedding that appears privacy-preserving on its own may reveal identity when information from multiple recordings of the same person is combined. We study this problem in clinical speech and introduce MoIRAGE (Multi-observation Identity-Resistant Adversarially Generated Embeddings), a framework for privacy-preserving representation learning under repeated observations. MoIRAGE separates an input embedding into clinically relevant content and an identity-related component, discards the original identity-related component, and replaces it with a synthetic alternative before generating the released embedding. The model is trained against identity-linkage attacks that jointly analyze multiple recordings, while clinical-consistency and diversity objectives preserve clinically relevant information and prevent representation collapse. At inference time, MoIRAGE does not require the participant's diagnosis label. We evaluate MoIRAGE on the UNMC clinical speech dataset using pooled held-out predictions from five participant-disjoint outer folds covering 137 participants, including 127 with clinical labels. MoIRAGE retains performance on the target clinical task, achieving an AUROC of 0.643 compared with 0.567 for the original embeddings. With a single recording (K=1), the strongest identity-linkage AUC* decreases from 0.778 to 0.528, close to chance. With 20 recordings (K=20), the strongest identity-linkage AUC* decreases from 0.963 for the original embeddings to 0.811 with MoIRAGE, while TPR at 1% FPR decreases from 0.474 to 0.036. Ablation experiments identify replacement of the identity-related component as the central mechanism underlying the privacy reduction, while a teacher-free variant shows similar behavior on Bridge2AI without speaker-teacher supervision. These results show that privacy cannot be adequately evaluated from single observations alone: identity leakage can increase as repeated measurements become available, and MoIRAGE explicitly addresses this multi-observation threat model while retaining performance on the target clinical task.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.