From Unique to Unified: Attribute-Centric Self-Supervised Representation Learning for Mixed-Type Tabular Data Clustering
Abstract
Mixed-type tabular data are common in scientific and real-world applications, yet learning reliable representations remains challenging because their attributes encode different relational semantics. Current self-supervised paradigms commonly derive supervision from inter-instance similarity, but such relations can be insufficient or unstable for heterogeneous tabular data. Moreover, numerical, nominal, and ordinal attributes encode fundamentally different notions of similarity. We therefore exploit inter-attribute dependencies as a complementary source of self supervision and follow a preserve-before-unify principle, in which type-specific semantics are first retained and subsequently integrated into a shared representation. Based on this principle, we propose Faithful Unification through Semantics-Preserving Encoding (FUSE), an attribute-centric self-supervised framework for faithfully unifying mixed attributes. Its Type-Faithful Attribute Formation (TAF) module aligns perturbation and recovery objectives with the relational properties of each attribute type, preserving their unique semantics before fusion. The Cross-Attribute Relational Unification (CRU) module further captures within-sample cross-attribute dependencies through target-exclusive contextual prediction, providing relational supervision for a unified representation. Comprehensive experiments across diverse mixed-type benchmarks demonstrate the efficiency and superiority of FUSE.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.