acceptodds
Under review as a conference paper at ICLR 2027

Let Androids Dream: From Visual Elements to Emergent Implications via Contextual Alignment

Abstract

Image metaphor understanding differs from general VQA: the implied meaning is rarely explicit in pixels and instead emerges from the compositional interaction of visual elements together with social intent and culture-specific background. We identify missing context – the gap between surface perception and the external knowledge required to infer emergent implications – as a failure mode of current multimodal large language models (MLLMs). We formalize Contextual Alignment as inference over a latent context: selecting background knowledge that is relevant to the image's composition and useful for inferring its implication. We then propose Let Androids Dream (LAD), a training-free framework that operationalizes contextual alignment through test-time computation: (1) Perception extracts compositional visual structure, (2) Search acquires missing context via perception-guided retrieval, and (3) Reasoning aligns compositions with context to derive emergent implications. On S100, a benchmark of 100 high-level implication questions in English and Chinese, LAD raises the lightweight GPT-4o-mini from 44% to 74% (English) and from 42% to 52% (Chinese) Multiple-Choice Question (MCQ) accuracy, on par with the closed-book GPT-5.4 in English, and from 2.98 to 4.02 and from 3.36 to 3.66 on Open-Style Questions (OSQ); with GPT-4o, LAD obtains the highest OSQ scores among directly comparable systems (4.14 and 4.26). LAD improves all four backbones we test, including the open-source Qwen2.5-VL-7B (+18 MCQ and +1.30 OSQ points in English), and its gains hold on the full II-Bench and CII-Bench. Adding other search tools raises MCQ accuracy in five of six cases but changes OSQ scores by to , whereas LAD's search stage improves both, suggesting that how context is acquired matters beyond access to search. Our OSQ judge is within one point of trimmed human means for 95.7% of answers, and LAD with web search also improves MMMU, SeedBench, and MMStar.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.