Generating to Generalize: Socratic Questioning for Compositional Continual VQA
Abstract
Continual visual question answering (VQA) requires models to retain previously learned knowledge while reasoning compositionally across different tasks, yet existing generative continual learning methods primarily emphasize knowledge retention and overlook cross-task composition. We address this limitation with a Mixture-of-Experts framework comprising dual-purpose experts that learn visual question answering (VQA) and visual question generation (VQG) within a shared multimodal large language model backbone. Each expert is implemented via low-rank adaptation to preserve task-specific knowledge. In contrast, knowledge distilled from prior experts provides complementary views to new experts, enabling compositional generalization without storing past raw experiences. A task-agnostic router selects experts at inference time. To more rigorously evaluate compositional inference in continual VQA, we introduce a human-annotated benchmark comprising real-world compositional questions. Extensive experiments on both standard and our compositional VQA benchmarks demonstrate that our approach retains knowledge and achieves stronger compositional generalization than prior generative continual learning methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.