acceptodds
Under review as a conference paper at ICLR 2027

Input-Aligned Grid-State Handoffs Enable Compositional Visual Execution

Abstract

In visual programs closed over grid states, learned operations can execute accurately in isolation yet fail when composed. We study which boundary representation converts available primitive transitions into execution when the program is supplied. Re-entering each intermediate result through the shared spatial color interface raises exact accuracy on unseen primitive pairs from 43.5% to 88.2% under matched training, with the advantage persisting on longer programs. With every learned weight fixed, the same handoff raises held-out-pair accuracy from 29.2% to 87.6%. Matched position-wise probes recover the intermediate grid from both executors, and grid-state re-entry makes that information directly consumable by the next operation. The advantage also transfers to grounded-object operations and programs recovered from demonstrations. For closed grid-to-grid visual programs, familiar operations compose when every intermediate result returns through the grid representation consumed by the next operation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.