acceptodds
Under review as a conference paper at ICLR 2027

R-CARE Reliability-Aware Concept Graph Conditioning and Evidence Verification for Visual Figurative Language Reasoning

Abstract

Visual figurative language understanding requires multimodal models to reason beyond the literal content of images and text by capturing implicit relations among visual cues, linguistic concepts, and figurative meanings. However, existing methods remain limited in structuring cross-modal figurative relations and in mitigating noise and semantic bias in automatically generated auxiliary knowledge. Moreover, explanations that are fluent and highly probable under the model may still be inconsistent with the multimodal evidence in the input. To address these challenges, we propose R-CARE (Reliability-aware Concept graph Adaptive Rhetorical Reasoning with Evidence Verification), a framework for visual figurative language reasoning. R-CARE first constructs a structured rhetorical concept graph that explicitly organizes visual entities, textual concepts, and potential figurative meanings. It then performs concept graph conditioned decoding, adaptively integrating graph structure and fine-grained semantic information based on the multimodal context while constraining their influence to mitigate interference from unreliable knowledge. Furthermore, R-CARE decouples label prediction from explanation verification. Given a fixed predicted label, it generates multiple candidate explanations, evaluates their consistency with the multimodal evidence using the model's internal multimodal representations, and selects the most evidence-consistent candidate subject to a generation-likelihood constraint. On the V-FLUTE benchmark, R-CARE achieves F1@0, F1@53, and F1@60 scores of 89.74, 78.14, and 60.85, respectively, surpassing MAPPER by 5.20, 8.45, and 9.13 percentage points. Code is available at: https://github.com/dahfs/R-CARE

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.