acceptodds
Under review as a conference paper at ICLR 2027

EAct: Coupling Structured Effect Space and Action Trajectories for Robot Control

Abstract

World-action models guide robot control by anticipating how a scene will change. However, similar visual transitions can arise from trajectories with different motion paths, orientation changes, and gripper timing. Learning to reconstruct future observations therefore does not ensure that the resulting representation captures how the change is carried out. We introduce EAct, which couples a structured effect space with a trajectory-supervised action representation. Its geometry-regularized effect space encourages transitions with similar visual changes to have nearby effect representations. Differences between frozen current and future visual features define the similarity reference, adding relational supervision beyond reconstruction. An effect-conditioned ActionVAE learns trajectory-level execution representations from demonstrations. Through effect-guided policy learning, the policy predicts an effect, decodes it into future visual features, and uses dense transition features to predict an action latent that conditions a flow-matching expert. EAct achieves average success rates of 99.0% across the four LIBERO task suites and 93.04% under RoboTwin clean evaluation, outperforming the compared baselines on both simulation benchmarks. On LIBERO-Object, an effect-only variant achieves comparable success to the original-latent baseline with 80% fewer Stage-2 updates. Code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.