RIPA: Rule Induction from Probabilistic Atoms for Closed-Loop Multimodal Inductive Reasoning
Abstract
Multimodal Large Language Models (MLLMs) have made substantial progress in visual understanding, yet remain less reliable in visual inductive reasoning and provide limited rule-level interpretability. Neuro-symbolic methods, by contrast, support explicit rule induction and traceable inference, but their integration with MLLMs for rule-guided prediction and image-specific rule verification remains insufficiently explored. In this paper, we propose RIPA, a closed-loop framework integrating neuro-symbolic rule induction with MLLM reasoning. RIPA first uses an MLLM to extract probabilistic atoms describing object attributes, relations, and overall scene structure from images. A Template-Driven Differentiable Rule Learner (TDRL) then induces sparse class-specific rules from these atoms and stores selected rules in a rule bank. During inference, TDRL scores candidate classes and retrieves their associated rules. These scores and rules, together with the image and probabilistic atoms, are fed back to the MLLM for prediction and image-specific rule verification. Finally, a consistency-aware module fuses the TDRL scores and MLLM outputs using rule agreement and conflict signals to produce the final prediction, thereby closing the reasoning loop. Experiments on two established object-centric visual reasoning benchmarks and our RAVEN-Custom benchmark show that RIPA improves over the neuro-symbolic MLLM baseline ILP-CoT by 10% on average. Moreover, RIPA exports explicit symbolic rules and rule-level evidence, making its predictions interpretable and traceable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.