acceptodds
Under review as a conference paper at ICLR 2027

StatePortal: Agentic Checkpoint Synthesis for Testing Deep States in Interactive Software

Abstract

Large language models have made from-scratch code generation for simple graphical user interfaces (GUIs) nearly free. For interactive software with long interaction histories, however, the development bottleneck has shifted from writing code to reaching the state under test: most target behaviors are not visible on the initial screen and become reachable only after a long sequence of interactions, and this reachability cost recurs on every commit. Existing GUI automation compounds the problem because it assumes a structured runtime representation (e.g., a DOM or accessibility tree) that self-rendered engines such as game clients and creative tools do not expose. We introduce StatePortal, an agent that synthesizes a checkpoint (a versionable, re-runnable configuration artifact) from a natural-language functional test case and read-only access to the target's source code. The checkpoint is restored through a per-application hook that a source-guided onboarding agent binds to an existing automation surface or synthesizes when none exists, once per application and without hand-written hook code, which decouples reaching a deep state from replaying the GUI path to it. To score checkpoints, we introduce hybrid programmatic-state and visual evaluation: acceptance is the conjunction of a programmatic-state oracle and a visual oracle, so the visual channel can only tighten the accept set, and we calibrate its residual error against a blinded cross-VLM reference. We release DeepStateBench, a benchmark of 61 tasks across 10 source-available interactive applications: 46 tasks on seven engine-backed games, our primary target, and 15 on three general-purpose GUI applications (JupyterLab, LibreOffice, and Blender). Across four models, a GUI agent seeded with synthesized checkpoints completes 10.4% of trials against 0% for the same agent unseeded under identical pixel input, while StatePortal, which also delivers the target action through the hook instead of by pixel input, completes 27.5%; 1361 replays over 112 checkpoints reproduce the original verdict at 98.90% with zero model calls; and the hybrid oracle detects 96.25% of controlled visual faults that leave every programmatic state field unchanged. We open-source our codes and benchmark at https://anonymous.4open.science/r/State_Portal-BB4E.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.