acceptodds
Under review as a conference paper at ICLR 2027

GenoFA: quantifying privacy leakage of genomic language models under dependent data

Abstract

Genomic language models (gLMs) enable large-scale training on personal genomic variants and advance personalized medicine. However, they also introduce unique privacy risks. Human genomes are inherently correlated across familial relationships due to Mendelian inheritance. Such correlation breaks the assumption of independence required by most differentially private algorithms, preventing straightforward adoption in the context of genomic privacy. We leverage Mendelian inheritance to create a family-aware genotype reconstruction attack, called GENOFA. gLMs can impute on non-training data points, but this should not count as privacy leakage as it occurs for seen and unseen data alike; we propose imputation-corrected reconstruction risk, a novel metric to remove the model’s imputation impact from privacy accounting. Our results demonstrate that familial correlations can amplify individual privacy leakage, motivating the need for family-level privacy protections. We therefore propose to treat the family as an atomic unit for privacy protection and apply client-level differential privacy to bound the privacy leakage of the entire family. We demonstrate that with a privacy budget of 6.0, client-DP improves the model utility by 22.9%, reducing the RMSE from 0.35 (group-DP) to 0.27. Together, these contributions provide a unified framework that identify, quantify, and mitigate privacy leakage in gLMs despite the inclusion of familially related individuals.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.