acceptodds
Under review as a conference paper at ICLR 2027

Learning Reusable Spaces of Feedback Policies

Abstract

Learning policies that combine diverse skills remains a challenge for robotics and embodied systems, particularly when a controller must reuse and adapt learned behaviors across new tasks. Existing approaches learn latent or task-conditioned policies and behavioral priors for diverse skills, but they do not explicitly learn a shared functional representation of the underlying feedback policies. We present a method for learning a low-dimensional space of feedback policies, where each policy is represented by a small coefficient vector, enabling identification, synthesis, interpolation, and analysis. Specifically, we learn neural basis functions from expert trajectories to span a low-dimensional subspace of feedback policies. The linear coefficient structure reduces policy identification to coefficient estimation, policy synthesis to low-dimensional optimization, and interpolation to direct operations in coefficient space. Across humanoid, quadruped, and planar-quadrotor control, the learned space identifies held-out policies from sparse demonstrations, supports synthesis and interpolation, and enables faster, more accurate adaptation than prior methods while maintaining competitive closed-loop performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.