acceptodds
Under review as a conference paper at ICLR 2027

Controllability of Pretrained Visual Representations under Start-Controlled Evaluation on Multi-Object Scenes

Abstract

Turning pretrained visual representations into behavior requires a procedure for sequential decision making, such as test-time search (CEM) via model-predictive control in the latent space or a goal-conditioned inverse-dynamics model (GC-IDM). We contribute a start-controlled evaluation protocol, an extended evaluation suite with progressively greater object counts, and a long-horizon predictive architecture for learning visual representations. Our start-controlled protocol shows that test-time search methods fail on 3D cube manipulation (0.0% success rate) and that GC-IDM still succeeds (68.9–84.9% success rate). Our extended evaluation suite scales the cube count up to four cubes, revealing that prior work's (LeWM and Sub-JEPA) representations become insufficient for downstream control under GC-IDM. The prior work's random-start evaluation protocol hides both issues, inflating LeWM's four-cube performance from 17.1% to 72.6%. In contrast, on four-cube tasks our proposed architecture learns representations that outperform those of LeWM and Sub-JEPA by up to 20.4 and 35.6 percentage points respectively. Our architecture predicts 50 environment steps of future latents in one pass and, through a structured attention mask, simultaneously learns inverse dynamics from its internal representations of consecutive states; a patch decoder maps each latent back to the encoder's patch representation, retaining object-level detail. Our representations are also more robust to colour distractors in one-cube pick-and-place. Ablations show that our method depends strongly on the training horizon, inverse-dynamics loss, and patch representation reconstruction loss, with the inverse-dynamics term alone worth a third to two fifths of control performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.