PROBABILISTIC MULTIMODAL REPRESENTATION LEARNING VIA PRODUCT-OF-EXPERTS AND QUANTUM TENSOR NETWORKS
Abstract
Multimodal representation learning must fuse heterogeneous signals while accounting for differences in modality reliability. Deterministic fusion typically does not represent this uncertainty, whereas generative fusion can be constrained by compact latent spaces and posterior collapse. We introduce QuTA (Quantum Tensor Autoencoder), which combines a Product-of-Experts posterior with a quantum Tree Tensor Network (TTN). The PoE forms a shared probabilistic latent from modality-specific Gaussian experts, and a 16-qubit TTN transforms it into nonlinear quantum features using only 75 circuit parameters. A classical skip and an objective combining task, reconstruction, and KL losses make the hybrid model trainable end-to-end. The architecture defines a deployment path from local simulator-based training to remote quantum-processor inference. We further evaluate the frozen circuit with noisy density-matrix simulation and observe stable predictive performance. Across four benchmarks, QuTA achieves the best result in the matched-protocol CH-SIMS comparison and remains competitive on MOSI, MELD, and MOSEI, while using substantially fewer trainable parameters than fine-tuned-BERT systems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.