acceptodds
Under review as a conference paper at ICLR 2027

DiReCT: Diagnosing and Repairing Context Vulnerabilities in Tabular Foundation Models

Abstract

Robustness evaluation is essential for tabular foundation models (TFMs) because real-world tabular data commonly contain imperfections. Unlike marginal metrics, which can overlook heterogeneous vulnerabilities across queries, query-conditional evaluation can reveal fine-grained failure patterns. For in-context TFMs, these vulnerabilities can arise from the context itself: large contexts can contain irrelevant, corrupted, or distributionally mismatched examples. Also, robustness studies should be interpretable for better diagnosis and treatment. However, current evaluations largely measure marginal degradation, often focus on query-side perturbations, and rarely combine interpretable diagnosis and actionable repair. We address these limitations with a training-free framework, DiReCT. We characterize queries using interpretable descriptors, measure their heterogeneous sensitivity to context perturbations, and statistically confirm the reliability of descriptor–perturbation vulnerabilities. We map confirmed vulnerabilities to deterministic context repairs using available operators, moving beyond diagnosis alone. Across diverse tasks and two state-of-the-art TFMs, we uncover recurring query-conditioned defects, and the repair results show that confirmed defects can translate into useful interventions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.