REBASE: Relational Basis Learning for Heterogeneous Multi-Source Tabular Data
Abstract
Multi-source learning for tabular data is fundamentally challenged by feature heterogeneity across distributed sources, arising from both structural heterogeneity (differences in feature dimensionality and composition) and semantic inconsistency (variation in feature definitions, naming, and units). This challenge is particularly pronounced in healthcare, where clinical variables are collected under institution-specific protocols. While such surface-level discrepancies exist, the underlying disease mechanisms are largely shared across populations, manifesting through complex interactions among clinical features. To address this challenge, we propose REBASE, a framework that learns reusable relational bases that can be shared across heterogeneous tabular sources. REBASE represents each sample as a graph, where nodes and edges are grounded in large language model embeddings, mapping heterogeneous features into a unified semantic space. Over this space, our model learns a set of relational bases, each capturing a distinct relational pattern. It aligns each sample graph to these bases using Fused Gromov-Wasserstein distance, which jointly accounts for semantic similarity and structural correspondence. This alignment produces sample-specific coordinates that compose predictions from the bases. Experiments on three real-world benchmarks demonstrate that REBASE consistently outperforms existing multi-source tabular pretraining methods, in both robust generalization and few-shot adaptation to unseen, data-scarce domains, with its learned relational bases exhibiting clinically meaningful patterns across domains. Code is available at https://anonymous.4open.science/r/REBASE-E3B2/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.