acceptodds
Under review as a conference paper at ICLR 2027

From Row-Level to Table-Level: Zero-Shot Private Attribute Inference in Tabular Data with Large Language Models

Abstract

Withholding sensitive columns is a common precaution when sharing tabular data, yet the remaining attributes may still support inference of the hidden values. We study this privacy risk under a zero-shot adversary that observes the released dataset and uses pretrained large language models (LLMs), without receiving ground-truth examples of the target attribute. The attack is evaluated under different levels of feature disclosure and with or without aggregate target proportions. Beyond direct LLM predictions for individual samples, our analysis asks whether statistical dependencies across the dataset can enable additional attribute recovery. We develop an evidence-based pseudo-context strategy that uses the LLM's scores to select pseudo-labeled samples for a tabular learner, deriving a separate context budget for each predicted class without ground-truth target labels. On demographic and health survey data, direct LLM inference is competitive with zero-shot baselines, and the table-level extension improves average accuracy and macro recall over direct prediction. These findings demonstrate how pretrained semantic knowledge can supply supervision for learning from an unlabeled dataset, exposing a source of privacy risk beyond independent sample-level inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.