acceptodds
Under review as a conference paper at ICLR 2027

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

Abstract

Compact robot manipulation policies need representations that capture how actions change observations. Large multimodal backbones and pixel-level world models provide semantic and predictive structure, but can devote substantial computation to details beyond those needed for control. We propose **SLIM** (**S**elf-supervised **L**atent **I**nteraction **M**odel), a compact policy that learns action-grounded predictive representations within the same backbone used for action generation. A Mixture-of-Transformers (MoT) jointly models observation latents and continuous action tokens. In the first stage, SLIM learns these representations through self-supervised masked trajectory prediction. It reconstructs actions from current and future observations and predicts future visual representations from current observations and actions, using an exponential moving average (EMA) target encoder to provide the prediction targets. The second stage builds on the visual representations learned in the first stage and trains the shared backbone for action generation with flow matching, requiring neither future observations nor explicit future generation at deployment. With 0.47B trainable policy parameters, SLIM achieves competitive performance on LIBERO and CALVIN and reaches 77.45% success on zero-shot LIBERO-Plus. It also outperforms and Fast-WAM in average task progress across five real-world tasks and four evaluation settings. In our H100 inference benchmark, SLIM achieves and speedups per inference call over and Fast-WAM, respectively, while reducing per-call FLOPs by over 76% relative to both baselines. These results support bidirectional latent prediction as an effective training approach for compact manipulation policies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.