Vulnerable Sets Need Not Yield Reliable Pointwise Rankings: Wrong-Label Causal Vulnerability in Two Tabular Foundation-Model Readouts
Abstract
Mechanistic studies of tabular foundation models can identify context examples that strongly influence prediction and use these signals to construct targeted poisoning attacks. Attack effectiveness, however, does not answer two stronger questions: whether high-influence positions are more vulnerable to wrong labels than matched-random positions under the same corruption budget, and whether the same importance score can rank this damage pointwise. We test both questions across two mechanistically characterized readouts, using deletion of the same rows to separate the harm of retaining wrong supervision from removing the samples themselves. Retaining wrong labels is more harmful than deletion for the high-influence sets selected by both readouts. For TabPFNv2’s attention-vote readout, the combined 20-dataset analysis shows a 2.62 accuracy-point excess [0.78, 4.65] for high-influence over class-matched random corruption, with clear cohort heterogeneity. In a separate 9-dataset pointwise analysis, the same pre-corruption importance shows weak correspondence with single-point damage: macro Spearman correlation is 0.069 for NLL damage and −0.011 for accuracy damage. TabICLv2’s prototype readout establishes neither stable set-level risk enrichment nor reliable pointwise ranking. Mechanism-derived importance can therefore identify more vulnerable sets without guaranteeing reliable pointwise risk ranking; the two properties require separate validation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.