PhysFlow: Physics-Aware Image Editing via Continuous Latent Thoughts
Abstract
Instruction-based image editing has achieved impressive progress in visual quality and prompt adherence, yet maintaining physical consistency remains challenging when edits involve complex dynamics such as refraction, deformation, or object interactions. Existing efforts either rely solely on textual reasoning, which struggles with fine-grained visual changes and lengthy token sequences, or resort to video generative models to synthesize intermediate transformation states, incurring prohibitive sampling overhead. In this paper, we present PhysFlow, an efficient continuous-thought framework that jointly internalizes textual editing planning and visual evolution within a compact latent space. By encoding the underlying transformation process as continuous latent thoughts, this way elegantly bypasses the limitations of both language-only reasoning and multi-frame rendering. Specifically, we train PhysFlow via a three-stage curriculum: first learning from explicit multi-step textual rationales and temporally ordered visual states, then progressively compressing them into continuous latent tokens, and finally steering the image editor with the learned latent editing dynamics. To enable this curriculum, we curate a high-quality dataset (namely PhysTrace) spanning diverse physical phenomena, annotated with hierarchical reasoning chains and dense visual trajectories. Extensive evaluations on PICABench and RISEBench demonstrate that our PhysFlow substantially outperforms competitive baselines in physical plausibility while introducing minor inference latency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.