ModalVerse: Multimodal Autoregression in Parallel Inference
Abstract
Multimodal autoregressive models commonly serialize different modalities into a single sequence, coupling modality-local causal progression with cross-modal interaction. We introduce ModalVerse, a framework that decouples these two operations to enable parallel yet cooperative multimodal inference. ModalVerse organizes modality sequences into parallel lanes that share an autoregressive backbone while maintaining separate causal histories. Within each layer, modality-local updates proceed concurrently, followed by cross-lane attention that exchanges information across modalities. Moreover, to sustain interaction when output lengths differ, continuation tokens keep completed lanes participating in cross-modal communication after visible generation ends. The framework supports two complementary settings: parallel multimodal input processing, where distributed evidence is integrated into a shared prediction, and concurrent multimodal output generation, where heterogeneous outputs advance independently while remaining coordinated. On MUSIC-AVQA, balanced parallel input lanes achieve up to faster decoder prefill and reduce peak memory by up to while maintaining comparable question-answering accuracy. On adapted A-OKVQA, ModalVerse improves joint text–speech accuracy from with independent parallel generation to , closely matching with concatenated generation. These results show that multimodal autoregression need not serialize modalities into a single sequence: modality-specific computation can proceed in parallel while preserving effective cross-modal interaction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.