MVSTab: Bridging Statistical and Semantic Subspaces for Tabular Representation Learning
Abstract
Self-supervised representation learning on tabular data predominantly relies on heuristic feature perturbations or ad-hoc column partitioning. However, these paradigms frequently suffer from a fundamental semantic-statistical gap: they optimize solely for empirical correlations while remaining blind to domain-level semantic structures. To bridge this divide, we propose MVSTab, a multi-view tabular representation learning framework that synergizes empirical statistical patterns with high-level semantic priors. Departing from heuristic masking, MVSTab introduces a dual-driven partitioning strategy that combines language model representations with correlation clustering to disentangle features into complementary semantic and statistical subspaces. To prevent representations from exploiting superficial numerical shortcuts, we design a semantic-aware hard negative generation mechanism. By selectively perturbing intra-group dependencies while preserving single-feature marginal plausibility, this mechanism compels the model to resolve nuanced logical contradictions rather than trivially detectable statistical outliers. Optimized via a tri-objective contrastive loss and aggregated through latent cross-attention, MVSTab produces dense, fixed-dimensional, and logically consistent embeddings. Extensive evaluations across diverse benchmarks show that MVSTab consistently outperforms competitive tree-based and deep learning baselines, especially in label-scarce regimes. Crucially, downstream analyses reveal that MVSTab uniquely excels at uncovering stealthy, logic-driven anomalies that evade conventional density- and reconstruction-based approaches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.