VisionCreator-P1: Closing the Loop for Long-Horizon Physical Visual Generation
Abstract
Long-horizon physical visual generation requires modeling how a scene should evolve under physical laws, not just how each frame looks. However, open-loop models lack explicit physical feedback: small local violations accumulate across steps and progressively derail multi-step trajectories. To address this, we present **VisionCreator-P1**, a closed-loop physics-informed agent that casts multi-image trajectory synthesis as a process of verified state transitions. At inference, the agent alternates milestone planning, next-state proposal, and transition-level physics reflection, committing a state only after iterative verification against the given query. For training, we address the credit assignment bottleneck that destabilizes long-horizon agentic reinforcement learning by decomposing optimization into short, prefix-conditioned blocks. Each block branches from a shared verified prefix and competes under a physics-aware hierarchical reward, thereby isolating local transition errors from accumulated drift and enabling stable policy learning. Experiments show that VisionCreator-P1 achieves state-of-the-art performance on physical accuracy and temporal coherence, establishing closed-loop local transition verification as a core paradigm for physically grounded, long-horizon generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.