From Seeing to Solving: Benchmarking and Diagnosing Symbolic Reasoning in Large Vision-Language Models
Abstract
Large vision-language models (LVLMs) remain unreliable on symbolic reasoning tasks, yet final-answer accuracy alone provides limited insight into whether failures arise from incorrect problem recovery or downstream symbolic solving. We curate the Multimodal Constraint Satisfaction Problem Benchmark (MMCSP-Bench), which provides matched structured-text, standard-image, and out-of-distribution (OOD) image views of Sudoku, Nonogram, and Graph Coloring, with all candidate solutions verified against the original constraints. We further propose the Declarative Constraint Satisfaction Problem Interface (DeCSP), a controlled symbolic interface that replaces unconstrained solver-program generation with predefined constraint-operator selection and deterministic execution. Across seven LVLMs, DeCSP achieves 96.24% mean valid-solution accuracy on structured text, compared with 8.19% for direct answering and 43.43% for PAL. However, DeCSP accuracy decreases to 42.57% on standard images and 28.19% on OOD images, revealing substantial challenges in multimodal symbolic reasoning. Controlled evaluations with fixed solvers, correct structures, and shared transcriptions disentangle visual-recovery failures from symbolic-generation failures, showing that recovery quality and final validity measure distinct dimensions of reasoning. Our stage-targeted training further shows that recovery improvements do not consistently transfer to direct answering, while program-training effects vary across models. Code and data are available at https://anonymous.4open.science/r/Seeing2Solving-37532.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.