Input-Aligned Grid-State Handoffs Enable Compositional Visual Execution
Abstract
In visual programs closed over grid states, learned operations can execute accurately in isolation yet fail when composed. We study which boundary representation converts available primitive transitions into execution when the program is supplied. Re-entering each intermediate result through the shared spatial color interface raises exact accuracy on unseen primitive pairs from 43.5% to 88.2% under matched training, with the advantage persisting on longer programs. With every learned weight fixed, the same handoff raises held-out-pair accuracy from 29.2% to 87.6%. Matched position-wise probes recover the intermediate grid from both executors, and grid-state re-entry makes that information directly consumable by the next operation. The advantage also transfers to grounded-object operations and programs recovered from demonstrations. For closed grid-to-grid visual programs, familiar operations compose when every intermediate result returns through the grid representation consumed by the next operation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.