Latent Rule Reasoning
Abstract
Rules are reusable abstract relations or constraints across different contexts. Inferring and applying rules to unseen cases across various tasks remains challenging for specialized models and Large Vision-Language Models (LVLMs). We propose **Latent Rule Reasoning (LRR)**, which integrates sequential rule induction and execution within an LVLM’s continuous latent space. LRR induces rule representations from visual context and applies them under query conditions to form target-state representations for answer generation. We further introduce **Latent Rule Self-Distillation (LRSD)**, in which the same LVLM uses complete visual observations available only during training to guide rule induction from inference-time inputs. Training first grounds rule representations and develops target prediction using rule-annotated examples, then incorporates cross-view rule consistency to jointly learn from examples with and without explicit rule annotations while retaining answer supervision. Our experiments reveal that LRR achieves an improvement of 10.70% on VisuLogic and a gain of 6.75% on SpatialEval, demonstrating the effectiveness of latent rule reasoning on various in-distribution, out-of-distribution, and general visual understanding tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.