acceptodds
Under review as a conference paper at ICLR 2027

MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning

Abstract

Understanding metaphors in images remains difficult for current AI systems. Multimodal Large Language Models (MLLMs) answer basic Visual Question Answering (VQA) questions well but often miss the cultural, emotional, and contextual implications of visual content, which require multi-hop reasoning, cultural context, and Theory of Mind (ToM). Reinforcement learning (RL) is a natural tool for eliciting such reasoning, but the target output of the task, an interpretation, cannot be verified automatically. We propose MetaphorStar, an end-to-end visual RL framework for image implication that addresses this by decomposing the reference interpretation of each image into a set of True-False propositions, each of which yields a binary, verifiable reward. The framework includes three components: the fine-grained dataset MetaphorQA, the visual RL method MetaphorGRPO, and the benchmark MetaphorBench. Trained with rewards computed only from these propositions and a format check, the MetaphorStar family (3B/7B/32B) raises the item-level True-False accuracy of the 7B base model from 28% to 70% and transfers to Multiple-Choice and Open-Style questions that are never seen during training; among 32 open and proprietary MLLMs, including the newest GPT-5.6-Sol, Gemini-3.1-Pro, and Qwen3.8-Max, MetaphorStar-32B obtains the highest True-False and Open-Style scores, while the strongest proprietary models remain ahead on Multiple-Choice questions. We further find that, in our setting, supervised warmup on distilled rationales lowers performance and collapses policy entropy, that the LLM judge commonly used for open-style evaluation ranks verbose SFT-trained models above more accurate RL models, and that the trained models also score above their base models on general visual reasoning benchmarks. We analyze parameter scaling, data scaling, reward weighting, and different backbones including Qwen3-VL and LLaVA. All models, data, and code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.