acceptodds
Under review as a conference paper at ICLR 2027

GeoAct: Reconciling Video and Action with Robot Geometry

Abstract

World action models generate a future video and an action chunk for a robot in one sampling process. The training loss does not tie the two outputs to the robot’s kinematics. As a result, the video and the action can describe different arm motions. We propose GEOACT, a sampler that uses robot geometry to reconcile the two outputs during generation. At selected denoising steps, a frozen geometric backbone and a trained joint readout estimate the arm configurations shown in the clean video prediction. Forward kinematics places this estimate and the predicted action in the robot base frame. A bounded proximal step moves the clean action toward the visual estimate. GEOACT reinjects the corrected action at the current noise level and evaluates the world action model again. The remaining video generation can then respond to the correction. On a shared backbone, GEOACT raises task success from 94.5% to 97.5% on LIBERO and from 81.8% to 89.0% on RoboTwin 2.0. Post-hoc correction with the same readout and the same budget reaches 95.6% and 83.9%. Ablations show that the gain requires a fresh reading of the updated video and a bound on each update. Two interventions per action chunk cost 1.22× the base latency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.