acceptodds
Under review as a conference paper at ICLR 2027

Joint-Force Matching: Contact Mechanics as Supervision for RL Critics

Abstract

Contact forces improve robot learning, and grasping demonstrations record them at every contact site and time step. In reinforcement learning (RL), the critic estimates the return of each action and guides the policy. Existing methods, however, predict each contact force or take it as an input, and they leave out how the forces load the hand's joints. Under elastic contact, these loads are the gradient of a single energy. Which structure in contact forces improves the critic's value estimates has remained unclear. We propose joint-force matching (JFM), which maps measured contact forces through the hand's Jacobians to joint loads and trains the negative joint-angle gradient of a scalar energy on the critic's representation to match them. JFM keeps force prediction and adds this gradient target. In simulated multi-finger grasping, JFM's critics have a lower return error than those of a force-prediction baseline at every pretraining seed, and JFM's offline-to-online policy succeeds on 33 of 158 held-out trajectories in closed loop against the baseline's 18. A free-vector head reaches a four to five times lower loss on the same joint loads, yet its critics have a higher mean return error than JFM's. We prove that JFM's energy targets the conservative component of frictional joint loads; the non-conservative remainder sets a floor on JFM's loss, and a free vector can also fit it. JFM keeps its advantage over the baseline when the energy gradient fed to the critic is zeroed at evaluation. JFM's gain in value estimates therefore comes from requiring the joint loads to be the gradient of one energy computed from the critic's representation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.