acceptodds
Under review as a conference paper at ICLR 2027

PolicyIK: High-Precision Inverse Kinematics for Robot Arms with Reinforcement Learning

Abstract

Inverse kinematics (IK) converts end-effector pose commands into joint motion, and many vision-language-action (VLA) policies rely on it to act on a given arm and to transfer across embodiments. Precise IK remains laborious: redundancy and singularities complicate the inverse mapping, execution depends on hand-tuned controllers, and training data on physical robots are costly to collect. We present PolicyIK, a goal-conditioned reinforcement learning policy that updates joint-position references from pose feedback and observes its accumulated reference relative to the measured joints, so the arm can rest on the goal rather than millimeters away. Two-scale pose errors with an error-dependent action gain let it correct millimeter residuals without slowing large motions, and rewards for progress and for staying within tolerance teach it to both reach and retain the target. In simulation, a single attempt keeps the mean errors over the final second within 1 mm and on 93.5% of held-out tasks, and a deployment protocol raises success to 98.9%. On physical hardware, PolicyIK achieves comparable or higher holding success than the evaluated classical IK controllers under matched execution budgets, with higher success rates under longer holding requirements. In simulated pick-and-place, it executes the end-effector commands of the VLA policy with 95.7% success. These results show that reinforcement learning can provide high-precision inverse kinematics for robot arms.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.