acceptodds
Under review as a conference paper at ICLR 2027

Analytic Policy Gradients for Contact-Rich Loco-Manipulation in Soft Robots

Abstract

Analytic policy gradients are widely considered impractical for contact-rich manipulation, where contact discontinuities induce bias and stiff or chaotic dynamics can produce high-variance estimates and exploding gradients over long horizons. Contrary to this view, we demonstrate contact-rich, soccer-inspired loco-manipulation in soft robots with up to 2,420 actuated springs, trained end-to-end with first-order gradients through simulation and control. On one commodity GPU, a single agent learned to locate, pursue, dribble, and kick the ball through the goal in 100 gradient steps—less than a minute of wall-clock time. In multi-agent settings, cooperative "passing" emerged strictly from task loss without explicit behavioral shaping, and in a competitive setting a defender tasked with preventing goals reduced its opponents' scoring rate by up to 60%. Gradient analysis reveals that sustained contact events inflate gradient energy by up to two orders of magnitude and leave per-window gradients largely misaligned, yet batch aggregation yields a reliable training signal. We also find that gradient energy is skewed away from the wide output layer that drives spring actuation, starving it of the sparse contact-mediated learning signal under plain stochastic gradient descent (SGD). With per-layer rescaling, SGD nearly matches Adam without per-coordinate adaptation. We release our simulator and task environment as a scalable, fully differentiable foundation for contact-rich, multi-agent learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.