PhysReflex: Counterfactual Cyclic Refinement for Physics-Consistent Video Generation
Abstract
Modern video generators achieve impressive visual fidelity, yet still struggle to faithfully capture the physical event specified by the prompt. This failure is particularly subtle when different event outcomes produce similar-looking motion, making current-state refinement potentially insufficient for distinguishing whether the intended dynamics are being formed. Existing methods address this challenge through physics-aware training, explicit physical priors, or inference-time guidance and refinement. However, training-free inference methods typically correct generation based on the current sampling state, without explicitly examining how different physical evolutions unfold from that state. We introduce PhysReflex, a training-free framework for counterfactual cyclic refinement. At selected sampling steps, PhysReflex briefly advances the current latent under the target condition, then traces it back under both the target and a counterfactual condition. The discrepancy between the returned states provides an evolution-aware signal for correcting the target sampling trajectory. We further adapt the number of refinement cycles according to changes in the predicted clean-video latent, avoiding unnecessary or excessive correction. Experiments on VideoPhy2 and PhyWorldBench demonstrate consistent improvements over baselines across comprehensive evaluations of physical plausibility, motion consistency, and semantic alignment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.