acceptodds
Under review as a conference paper at ICLR 2027

COMB: A Factorized Object-Centric World Model, and What Its Structure Buys a Policy

Abstract

Object-centric world models decompose scenes into entities, but each entity is typically represented by a single unstructured embedding. We study whether policy learning also benefits from a second decomposition, in which each entity is represented by named attributes such as position and appearance. This factorization can, in principle, support transfer to attribute combinations that were never observed during training. We introduce COMB, an object-centric world model with explicitly factorized per-entity latents, learned without attribute supervision and trained for online model-based RL in imagination. It achieves the highest or tied-highest return on eight 2D and 3D manipulation tasks and leads the baselines on held-out attribute combinations. To separate representation format from information content, we freeze COMB and insert an invertible map between its representation and the policy heads. The map preserves all information, changes only the coordinates in which it is expressed, and can be varied continuously. On 2D and 3D tasks with held-out position-colour combinations, entangling the attributes reduces return by 31-37% at a fixed budget, with larger losses at higher doses, and reduces held-out return by up to 30% at matched in-distribution skill. Linear scrambling produces no resolvable loss. These results suggest that, once entities are separated, policy learning benefits from making task-relevant attributes linearly readable; the effect therefore depends on both the representation and the head that reads it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.