A Universal Schema for Relational Foundation Models
Abstract
Traditional relational deep learning models use a database's relational structure to predict properties of its entities for a given task specification, but are tied to the schema and task on which they are trained. Recent work on relational foundation models aims to remove this constraint by pretraining a single model across databases and transferring it to unseen ones. Existing approaches address schema variation through the model architecture. We propose a complementary, representation-level approach: an information-preserving mapping that represents any database and task using a fixed universal schema, moving schema agnosticism from the model architecture to the input data. Any existing relational deep learning model can then be trained across databases and transferred to new ones. On RelBench classification and regression tasks, we show that key relational foundation-model capabilities can emerge from this representation alone. A standard GraphSAGE trained across databases achieves zero-shot transfer to an unseen database under the Relational Transformer protocol, while transferred models require substantially fewer examples to reach a given accuracy when fine-tuned on a new database. Overall, our results suggest that relational foundation-model capabilities need not be encoded entirely in specialized architectures: they can emerge from the representation itself, enabling a modular interface between relational data and learning architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.