acceptodds
Under review as a conference paper at ICLR 2027

VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement

Abstract

Video generative models are increasingly used to simulate physical interactions and support downstream decisions, yet their evaluators still tend to return scalar scores or free-text flaw lists. Such outputs make it difficult to tell which prompt requirement was satisfied, what evidence was resolved, and where physics laws were silently disobeyed. We present VeriPhy, a trace-first critic that evaluates generated video through executable, claim-bound evidence. VeriPhy decomposes the prompt into span-bound claims and a typed check plan. Then it composes a tool chain for each plan, and performs execution that records each check as measured, abstained, or errored. These records are finally assembled into verdicts that determine the physics alignment of each video. This trace supports diagnosis, learning, and refinement. Such a planning architecture also allows the introduction of in context learning capabilities, training free by prompt injection. VeriPhy demonstrated strong learning performance from human annotations, raising the flaws found on held-out clips significantly, and correlates more strongly with human semantic-adherence ratings than VideoScore2. Finally, VeriPhy can be directly wired back into the generation process for trace-guided prompt rewrite, and demonstrated better generation performance training free. Together, these results position executable evidence traces as an interface for auditing, improving, and using physical-video evaluators.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.