acceptodds
Under review as a conference paper at ICLR 2027

Synthetic Data Shield: Obfuscating Contexts for In-Context Tabular Prediction

Abstract

Tabular foundation models are becoming a popular approach to tabular classification based on in-context prediction. To classify a new entry, the model receives the labeled training set and the new row, then predicts its label in a single inference pass without task-specific training on the target dataset. A major privacy concern of this paradigm is that when a third party performs classification, the training data used as context must also be shared with that party. We propose **Synthetic Data Shield** (SDS), a method that does not require sending real training data to a third party. SDS creates a local synthetic training context and sends only this synthetic data, keeping the original training records hidden. The foundation model remains unchanged, while accuracy stays close to that of the original data. We evaluate our method through three privacy axes: membership inference, near-record exposure, and attribute inference. Our rigorous analysis shows that prediction utility, membership leakage, near-record exposure, and attribute inference relate to different statistical properties of the synthetic data generator, allowing us to control membership and near-record exposure separately while attribute inference tracks model accuracy in our measurements. Membership attacks stay close to chance, and synthetic records sit nearly as far from the training rows as from held-out rows of the same population. Attribute inference differs because statistical relations useful for prediction can also help predict a hidden attribute. We further show that adding noise can reduce this leakage but also causes a small loss in model accuracy. We test SDS on 71 tabular datasets and across tabular foundation models such as TabPFN-2.5, TabICL, LimiX, TabFM, and TabuLa-8B. We use the same synthetic contexts across these models (TabuLa-8B reads a 32-shot subset of each), with the mean accuracy loss staying within 2.5 percentage points for the tested transfer models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.