Visual Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models
Abstract
We propose ILP-CoT, a method that bridges Inductive Logic Programming (ILP) and Multimodal Large Language Models (MLLMs) for visual logical rule induction, which involves both discovering logical facts and inducing logical rules from a small number of unstructured visual inputs. This task still remains challenging when solely relying on ILP, due to the requirement of specified background knowledge and high computational cost, or MLLMs, due to the appearance of perceptual hallucinations. Based on the methodology that let MLLM be the input perceptor and rule structure proposer, meanwhile let ILP system serve as the rule inducer and symbolic verifier, our approach automatically builds ILP tasks with pruned search spaces, and utilizes ILP system to output rules built upon rectified logical facts and formal inductive reasoning. Its effectiveness is verified through challenging logical induction benchmarks, as well as a potential application of our approach, namely text-to-image customized generation with rule induction. Our code and data will be open-sourced upon formal release.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.