acceptodds
Under review as a conference paper at ICLR 2027

OrigamiWorldBench: Benchmarking Visual Instruction-Conditioned Next-State Generation

Abstract

Recent advances in image editing models have largely centered on natural-language instructions. In real-world diagrams and manuals, however, actions are often specified visually through arrows, lines, and symbols and may induce state transitions of the target object. Generating the resulting state from such visual instructions therefore requires models both to interpret visually specified actions and to faithfully realize their geometric and structural consequences. To study this capability, we introduce OrigamiWorldBench, a benchmark for next-state generation conditioned on visually specified actions. Origami is well suited to this setting because origami operations can be visually specified and their outcomes obey clear geometric and structural constraints. Given a current-state image overlaid with a visual prompt and an auxiliary text instruction describing the operation, models directly generate the corresponding one-step-ahead state image. To evaluate geometric and structural fidelity, which is difficult to capture with conventional image-similarity metrics, we introduce an aspect-wise VLM-as-a-Judge protocol. Using OrigamiWorldBench, we evaluate a set of recent image generation and editing models. To quantify the gap between current models and human performance under the same visual instructions, we additionally compare model predictions with origami states physically folded by human participants. We further validate the proposed aspect-wise VLM evaluation against human judgments and conduct an input-cue analysis. We plan to release the benchmark dataset and code upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.