Learning When Structure-Based Semantics Help Tabular Imputation
Abstract
Pretrained language models provide rich semantic priors for tabular imputation, but their utility is inherently conditional: general semantic associations may conflict with dataset-specific dependencies or become redundant when the observed values already provide sufficient evidence. Existing language-enhanced imputers therefore face a fundamental mismatch between general-purpose semantic priors and the instance-specific evidence required for each missing entry. To address this mismatch, we propose SemImputer, a structure-based semantic representation imputation framework that learns when semantic priors provide complementary evidence for each missing-value prediction. Specifically, SemImputer first aligns semantic representations with tabular structures through structure-semantic representation alignment, then encodes the missingness context for each target through context-aware missing pattern encoder, and finally learns whether semantic information improves each prediction through counterfactual utility-guided fusion for selective semantic injection. Across 10 real-world datasets, SemImputer achieves state-of-the-art performance, reducing RMSE by 5.4% and improving R² by 16.1% on average over the strongest baseline, while remaining robust across heterogeneous missingness mechanisms without mechanism-specific retraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.