acceptodds
Under review as a conference paper at ICLR 2027

Latent Skill Acquisition and Composition for Contact-Rich Manipulation

Abstract

Human-level generalization in manipulation requires composing known behaviors to reach novel goals in scenarios never seen in training. The challenge sharpens in contact-rich settings, where behaviors must be recomposed at the granularity of contact rather than of a semantic subtask. Without relying on per-task demonstrations, a promising approach is to leverage test-time compute to perform physical reasoning that searches for novel solutions and improves them incrementally during deployment. To effectively perform search in the large space of motions, we propose SkillReasoner, which learns temporal abstractions of actions (i.e., skills) from unstructured offline interactions, and then reasons over and recomposes them at test time. Crucial to the formulation is a latent skill space learned with varying temporal horizons, so that composition can happen at varying granularity, as required by different stages of the task, while minimizing the number of decisions. To realize a novel goal at test time, the robot searches over latent skill sequences in its imagination, decodes a chosen skill into low-level actions, and replans in closed loop, up to 30 Hz. We demonstrate on a series of contact-rich real-world and simulated tasks that the robot chains behaviors to achieve novel goals that require far longer behaviors than those seen in training, and adapts zero-shot to unseen perturbations, changing dynamics, and constraints, all trained purely from unstructured play data without task-specific demonstrations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.