acceptodds
Under review as a conference paper at ICLR 2027

Vec2DNA: Feature-Consistent Genomic Embedding Inversion Attacks and Task-Aligned Defense

Abstract

Genomic foundation models are increasingly produced through mean-pooled embeddings, yet it remains comparatively underexplored when it comes to what such vectors leak about the underlying DNA from a privacy perspective. We introduce Vec2DNA, a black-box embedding inversion framework that constructs single-nucleotide-perturbed anchor–variant pairs and exploits their known correspondence through a feature-level consistency objective. Rather than treating these variants as conventional data augmentation, Vec2DNA uses their pairwise structure to preserve internal representations at unchanged target positions while localizing the effect of sequence perturbations. Across four genomic embedding models on hg38 20-mer dataset, Vec2DNA consistently outperforms alternative inversion methods in reconstruction quality. In particular, it achieves up to a 14.5% relative improvement in Levenshtein similarity over standard inversion baselines. Our analysis shows that the structured correspondence between anchor–variant pairs provides the key additional signal through feature-level consistency. Beyond sequence reconstruction, we further show that genomic embeddings retain sufficient fine-grained information to reliably distinguish single-nucleotide variants and recover allele identity well above chance. Furthermore, we propose Task-Aligned Genomic Quantization (TAG-Q), a utility-aware defense that selectively preserves task-relevant embedding directions while suppressing information useful for inversion. TAG-Q substantially reduces inversion performance while retaining near-clean downstream utility on DNABERT-2, and achieves a similarly favorable privacy–utility trade-off on GENA-LM. Together, our results highlight the privacy risks of genomic embeddings and demonstrate complementary attack and defense mechanisms for controlling fine-grained genomic information leakage.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.