Can AI Speak the Language of Origami?
Abstract
Building AI systems that can plan, act, and create in the physical world requires more than pattern recognition. Such systems must reason about the generative mechanisms and constraints that govern physical processes, using structured representations that connect observations, actions, and their effects. Yet existing benchmarks often study these capabilities in isolation, focusing either on visual recognition or on abstract symbolic and programmatic reasoning. Origami offers a natural testbed for integrating these abilities: constructing shapes through sequences of folds requires visual perception, reasoning about geometric and physical constraints, and multi-step planning, while remaining sufficiently structured for systematic evaluation. We introduce OrigamiBench, a benchmark for evaluating programmatic understanding of the mechanisms underlying origami synthesis and the ability to induce such mechanisms in a high-level language of physically grounded fold actions. Experiments with modern vision-language models show that increasing model scale alone does not reliably improve reasoning about physical transformations. Moreover, models struggle to ground programmatic representations in visual observations, indicating that their visual and linguistic representations remain only weakly integrated.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.