CHAT: Automated, Constrained and Hard-Label Attacks on Tabular Models
Abstract
Evaluating tabular models' adversarial robustness is particularly challenging, as input perturbations must satisfy domain constraints (i.e., semantic relationships between features) to reflect realistic samples. Most existing attacks rely on manually created constraints which are inherently limited. Moreover, as they rely on model gradients or output scores, existing attacks are inadequate for practical settings where only hard labels are available. We address the first challenge by leveraging a large language model (LLM) to automatically mine rich, high-quality constraints. Our evaluation of constraints derived for three datasets from TabularBench shows that LLM-generated constraints almost fully capture manual constraints, cover additional critical constraints not found manually, and score highly on quality assessment by domain experts. We address the second challenge by proposing the first hard-label practical black-box attack on tabular models. By interweaving a misclassification-aware projection method for satisfying constraints and a leading hard-label attack from vision (CGBA), our attack, termed CHAT, achieves higher constrained attack success rates than the state-of-the-art white-box attack (CAA) on three datasets and eight models each, both defended and undefended.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.