FrameTrust: Outcome Agreement for Reliable Video Imagination
Abstract
As video models evolve into implicit world models, autonomous agents can "imagine before acting" - yet visual plausibility does not guarantee predictive reliability, risking irreversible errors. To bridge this gap, we introduce FrameTrust, a training-free visual self-consistency framework that isolates true outcome disagreement from mere path diversity by mapping diverse rollouts to discrete terminal states via task readers. For stochastic environments, a simulator-relative entropy contrast effectively disentangles genuine world randomness from model hallucinations. To mitigate extreme video sampling costs, an intermediate diffusion readout screens divergent futures before all generation steps are spent. These signals drive an intent gate that dynamically routes the agent toward execution, replanning, or a conservative fallback. Across spatial, physical, memory, and interface environments, FrameTrust significantly improves error ranking and reduces denoising work through early screening. Crucially, our gate reduces irreversible GUI execution errors by 59-64% relative to blind imagination without compromising task success, establishing a rigorous foundation for safe, imagination-driven autonomous action.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.