RISE: Co-Synthesizing Visual Games and Observation-Grounded Demonstrations
Abstract
Long-horizon visual tasks can require agents to learn how actions affect an environment and apply that knowledge in later decisions. Demonstrations generated by reading a simulator's internal rules can skip the interactions through which an agent would learn how the environment works. We introduce RISE (Rule Inference from Synthesized Environments), a framework that jointly generates executable visual games, connected tasks, and demonstrations based on public observations. Tasks are organized into multi-level sequences that share the same underlying game mechanics. Reusable symbolic solvers collect demonstrations through public images and interaction feedback, without per-step LLM teacher calls. For compositional puzzles, RISE uses observations from early interactions to construct later goals. Their verified solution paths include actions not yet tried in a given state, with outcomes predicted from the recorded observations. Task constructors check solvability, while replay checks link structured evidence records to their source observations and verify executed outcomes. Construction audits across nine game types test execution, evidence provenance, and export consistency. In a separate sixteen-game development study using an earlier generator, synthesized supervision improves task execution, and a revised data recipe outperforms the standard recipe at a similar target-token budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.