acceptodds
Under review as a conference paper at ICLR 2027

Certified Context Integrity for In-Context Predictors

Abstract

Modern in-context predictors treat a support set as part of the inference input: tabular foundation models condition on support tables, and large language models condition on demonstrations. This opens a poisoning surface that needs no retraining, because whoever controls the context can change predictions without touching the model. Against four tabular foundation models and 15-NN on six real tables, a median of one to eight inserted rows flips an undefended prediction, and for TabICL a single row suffices on half the tables. We introduce *certified context partitioning*, a retraining-free, black-box wrapper that issues a deterministic certificate for each query. Each support row is assigned to a shard by a cryptographic hash of its own features, the frozen predictor runs once on every realised shard, and the prediction is the plurality of the shard votes. An insertion, deletion or relabelling can change only one shard, so the vote gap certifies how many such operations the prediction survives; feature edits receive half that radius. Across 20 OpenML/UCI tables, 12 base learners, shard counts up to 256 and three seeds, a median certified radius of 7 rows (m = 16) costs a median 1.0 accuracy point. With every configuration chosen on held-out validation queries, partitioned in-context prediction beats certified k-NN, the closest retraining-free baseline, by 0.67, 1.00 and 0.94 points of certified accuracy at radii 1, 2 and 4 (winning on 14, 14 and 13 of 20 tables), while the same certificate over classical learners does not; under arbitrary edits the two are level unless rows carry trusted identifiers. Two white-box adaptive attacks that know the model, hash and partition produce no certificate violation in 3,600 attacks on 1,800 defended queries, and for TabICL, TabPFN v2.6 and Mitra v2 the median successful insertion attack needs exactly r + 1 rows, so the bound is attained. Content addressing costs no measurable accuracy against a balanced random partition, and partitioning cuts peak GPU memory 12–29× at 20,000 support rows. A learning-curve law predicts certified accuracy to a median error of 0.6 points and picks the best shard count outright in 60–74% of cells. Changing only the partitioner and backend certifies Qwen3.5-2B and Qwen3.5-4B on SST-2 at 89.0% and 88.0% against any eight inserted, deleted or relabelled demonstrations out of 512. Context supplied at inference time can therefore be treated as a certifiable security boundary, with deterministic guarantees and no retraining.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.