acceptodds
Under review as a conference paper at ICLR 2027

Don't Call It Privacy Until You Pick a Metric: Joint Auditing of Fine-Tuning and Privacy Metrics in Clinical LLMs

Abstract

Fine-tuning open-weight LLMs on clinical data improves task performance, but it can also leak sensitive information from the training data. There is no single agreed way to measure this leakage. This raises a practical question: for a given privacy budget, does full-parameter or parameter-efficient (QLoRA) fine-tuning give better utility? Answering this question is complicated by the fact that privacy itself has no single agreed-upon measurement, and three attack families have each been invoked in the literature to justify when a clinical LLM is safe, namely 1) verbatim extraction, 2) prompt-based identity probing, and 3) membership inference. We posit that these problems are not separable. The verdict on which fine-tuning regime supports the greatest amount of privacy depends on which attack one is concerned about. In this paper, we provide the first benchmark that varies both objectives jointly. We fine-tune six open-weight models on four clinical tasks. Each model is trained under regimes: full-parameter and QLoRA, each with two loss variants. We evaluate every resulting model with all three attack families and with task-specific utility metrics. We additionally study differentially private fine-tuning on a focused slice of the grid as a case study of how a privacy-preserving training method behaves with respect to various privacy attacks. In total, our study spans over 250 fine-tuned model configurations and more than 1,000 evaluation runs, including multi-seed replication, hyperparameter sensitivity sweeps, and decoding-controlled ablations. Our investigation yields two primary findings: (i) Privacy metrics can disagree, and the way they disagree is itself task-dependent e.g., verbatim extraction risk and membership-inference risk show negative correlations on one of our four tasks, with weaker or inconsistent relationships on the others and (ii) the fine-tuning regime interacts with these privacy metrics in a nontrivial, task-dependent manner. In other words, a fine-tuning regime may reduce privacy risk under one metric while increasing it under another. Privacy outcomes therefore depend on the specific combination of fine-tuning regime, attack, and clinical task. We release all code for a public benchmark that supports systematic evaluation of both privacy and utility in clinical LLM fine-tuning across multiple attacks and tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.