acceptodds
Under review as a conference paper at ICLR 2027

Think in Graphs: Disentangled Graph-Centric Multimodal Reasoning via Reinforcement Learning

Abstract

Reinforcement learning has significantly improved reasoning capabilities in large language models, yet extending these advances to vision-language models remains challenging due to insufficient modeling of spatial and logical relations among visual entities. Existing methods rely on isolated visual entities and final-answer supervision, which causes reasoning to drift toward language priors rather than visual relational structures. We propose GRATS, a two-stage multimodal reasoning framework that consists of constructing relational graphs from images and performing graph-centric chain-of-thought reasoning over the generated graphs. We further introduce a novel RL method with stage-specific credit assignment to optimize the model's graph construction and graph reasoning capabilities simultaneously. Experiments on seven multimodal reasoning benchmarks show that GRATS consistently improves performance on both Physical Perception and Diagram & Quantitative Reasoning tasks. Code is available at: https://anonymous.4open.science/r/GRATS.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.