acceptodds
Under review as a conference paper at ICLR 2027

World-Model Predictive Displays for Teleoperation

Abstract

Network latency degrades precision in teleoperation, and becomes critical in domains such as remote surgery. Predictive displays show an estimate of the present instead of the delayed camera feed, but learned image-space predictors fix the prediction horizon at training time and are valid only at that delay. We instead build a predictive display around an interactive world model, reaching the present by repeated single-step prediction from the most recent received frame. Building on DINO-WM, we predict in the feature space of a frozen encoder and decode only the displayed frame. A display thread performs one predictor and one decoder evaluation per frame regardless of the delay, while an asynchronous grounding thread re-anchors the latent context to newly arrived observations, so the compensated delay becomes a runtime quantity and jitter requires no retraining. In a within-subjects study with 22 participants on a simulated Push-T task at round-trip delays of 500–1000 ms, spanning the range where fine manipulation breaks down, compensation shortened completion time on successful trials by 23.5% and suppressed the move-and-wait strategy latency induces. Participants rated the display as more responsive but less authentic, and authenticity tracked rollout fidelity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.