acceptodds
Under review as a conference paper at ICLR 2027

Goal-Conditioned Multi-Object Manipulation via Compositional Actions

Abstract

Goal-conditioned reinforcement learning (GCRL) for multi-object manipulation from offline images promises robust, scalable, and generalizable control, but existing generative subgoal methods struggle in test and out-of-distribution settings where combinations of object state-to-goal trajectories lie far from the training distribution. We present COMPositional ACTions (COMPACT), a compositional GCRL framework that learns generalizable manipulation policies end-to-end. COMPACT combines an object-centric encoder that discovers objects via feature connectivity and factorizes them into appearance and pose representations, with a compositional policy that aligns state and goal objects by appearance and selects the least displaced object for action prediction. The selected state-goal object pair, together with a learned pose difference embedding, conditions a goal-conditioned policy trained with implicit Q-learning. In multi-object manipulation environments based on OGBench and Push-T, COMPACT improves task success and sample efficiency over state-of-the-art generative subgoal and flat baselines. Importantly, COMPACT is the only method that achieves non-trivial zero-shot performance in settings with an increasing number of objects. Our results demonstrate that object discovery and representation, appearance-based alignment, and compositional action selection provide an effective foundation for data-efficient, generalizable multi-object manipulation from offline visual data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.