acceptodds
Under review as a conference paper at ICLR 2027

When Worlds Meet: Benchmarking Implicit Causal Reasoning with IRS-Bench

Abstract

An image records a partial snapshot of a local world. Combining foreground and background snapshots therefore requires more than harmonizing pixels: the merged scene must realize the state changes and scene effects licensed by their interaction. Existing benchmarks for image composition primarily measure perceptual quality, identity preservation, and spatial alignment, leaving this merge-and-reaction problem untested. We formulate it as image-conditioned implicit causal reasoning and introduce IRS-Bench. Each instance provides a foreground image F, a background image B, and an instruction q specifying the entities and their placement while withholding the governing relation and its consequence. The source images fix entity identities; post-interaction states and scene effects must be inferred from visual evidence. To accept valid alternatives, IRS-Bench distinguishes deterministic interactions with one admissible state from intentional interactions with an admissible set. The benchmark contains 1,000 pairs across eleven leaf tasks and five causal mechanisms: physical, biological, chemical, institutional, and functional. Each item includes a governing rule, admissible outcomes, and weighted binary visual checks; 172 items also provide photographic post-interaction references. Under the ME-Ruler score, the strongest non-ReactionGraph baseline obtains 61.1 and 71.6 under GPT-5.6 SOL and Qwen-3.8-27B, respectively, across 18 systems. A training-free ReactionGraph reference pipeline obtains 77.2 and 88.8 under the same judges. These results show that perceptual coherence and causal faithfulness are distinct capabilities and that models must carry interaction constraints across visual contexts rather than merely produce plausible composites.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.