acceptodds
Under review as a conference paper at ICLR 2027

Learning a Synergy Basis from Reward Alone

Abstract

Musculoskeletal control involves action spaces where muscles far outnumber the mechanical degrees of freedom they drive, creating a redundancy that challenges standard reinforcement learning (RL). A natural remedy drawn from motor neuroscience is to act through a smaller set of muscle synergies. Most existing synergy-based RL methods obtain their basis from a separate discovery phase, a dynamics model, or physiology rather than from the task reward. We introduce Plexus, a latent policy learning method that constrains actions to non-negative combinations of synergy vectors and makes the basis a free parameter of the actor. Crucially, this basis is learned end-to-end, from reward alone, at the same time as the actor. Because the decoder is learned and otherwise unconstrained, it is generically non-orthogonal and distorts volume, so latent and ambient entropy no longer represent the same quantity. This makes the basis lose rank and the policy steer a subspace narrower than its latent size advertises. We remedy this by measuring entropy on the induced action manifold, through a Gram determinant correction to the log-density. We show that when the decoder is jointly learned, letting this quantity flow through the decoder keeps the basis well conditioned. Across six MyoSuite tasks on two musculoskeletal models, Plexus matches or exceeds full-dimensional baselines on every task and is competitive with synergy-based methods at matched latent width, while yielding a single state-independent basis that can additionally be initialized or constrained with biomechanical priors, or inspected for interpretability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.