FORCE: Factorized Object Relations for Counterfactual Editing in Video World Models
Abstract
World models predict how scenes evolve, but they are difficult to interrogate: given a model that forecasts a video, one cannot easily ask what would have hap- pened had a particular interaction between two objects been different. We present Factorized Object Relations for Counterfactual Editing (FORCE), a world model that infers pairwise interactions from raw video without supervision and repre- sents each as an explicit latent per object pair, so that a relation can be removed, changed, transferred or introduced directly. Objects are extracted by slot attention over frozen self-supervised features; a relation encoder infers each pair latent from a short observation window, and an interaction head turns it into a force. The latent is held fixed across the rollout, so an edit persists rather than being recomputed, and every edit acts on the model’s own quantities with no mapping to simulator parameters. We evaluate in a simulated environment where the true outcome of an edit is available, scoring the change each edit induces against the change the simulator produces under the same operation. On held-out clips, FORCE’s ed- its made from video agree with the simulator’s response at 0.19 to 0.37 on our effect-based score; trained and scored on exact object positions instead, the same architecture reaches 0.57 to 0.85. The gap therefore lies not in the learned rela- tional dynamics but in perception: what bounds an edit is how precisely objects can be located from video, and the precision editing needs lies just beyond what current object-centric perception delivers. Closing that gap is a concrete target for making counterfactual world models work from video.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.