Symbolic Reasoning Groundings for Latent Visual Language Models
Abstract
Neuro-symbolic learning combines flexible neural perception with structured, executable reasoning. Yet systems that place a predefined executor on the inference path bind their answers to a fixed ontology, limiting integration with open-ended vision–language models. We introduce neuro-symbolic supervision for latent reasoning: a framework that trains over continuous reasoning states by grounding them as shared entities and executable predicates. Symbolic losses thereby shape the latent trajectory from which the original VLM backbone generates its answer. This brings compositional supervision into latent reasoning while allowing the decoder to answer beyond the symbolic vocabulary. We instantiate the framework in through local symbolic grounding, compositional symbolic grounding, and reinforcement through discrete symbolic execution. The final stage pairs each sampled text with one symbolic scene and optimizes both using answer and execution rewards. The symbolic path is used only during training. With four latent steps, reaches a 72.53% mean across six benchmarks, achieving state-of-the-art performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.