Transfer Learning for Multi-Channel HDP-LDA Topic Models Under Strict Data Separation Constraints
Abstract
Latent Dirichlet Allocation (LDA) models are increasingly used beyond text, including for discrete tabular data in which samples are treated as documents and multinomial variables as observation channels. Such models have found renewed success in the biomedical domain, where datasets are usually small and targeted to a specific health condition. Often, such models can be improved by including data from external sources and transfer learning is the canonical approach in such situations. However, patient data is extremely sensitive and hosted on isolated sandbox environments whose strict governance rules forbid exporting individual-level data to other systems and require manual approval of every data export request. These restrictions make traditional transfer and federated learning methods for LDA models inapplicable. In this paper, we therefore propose new methods for transfer learning for multi-channel LDA models in the presence of strict data-separation constraints and where no iterative communication between data hosting systems is allowed. We tackle this challenge by extracting sufficient statistics from a model fit on the source dataset and using them as frozen, weighted priors when modeling the target dataset, thus requiring a single data transfer between systems. To approximate the topics that would be obtained by combining the datasets, we tune the weight of the frozen source summaries to match their empirical covariance with the nominal covariance they induce in the target model. For target-topic estimation, we introduce exact auxiliary-variable Gibbs updates that infer the transfer strength and the global concentration of a Hierarchical Dirichlet Process (HDP) prior under the asymmetric concentrations induced by the frozen source summaries. Controlled synthetic experiments across a range of sample sizes and domain shifts characterize the methods' behavior, while a case study on the Cleveland heart disease dataset provides a real-world demonstration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.