KINO: 3D Keypoint-Mediated State-Image Reasoning for Robot Pose and Joint Estimation
Abstract
Single-view robot pose and joint estimation requires recovering an articulated state whose pose, joints, and visible parts must be spatially consistent under robot kinematics. Existing methods establish this consistency either implicitly through direct regression or explicitly through render-and-compare refinement. The former leaves the relation between image observations and 3D articulated states implicit, while the latter makes consistency checking explicit in the image plane but still requires the refinement model to infer 3D state corrections from 2D discrepancies. In this paper, we propose KINO, which formulates robot state estimation as latent geometric reasoning between image-induced and state-induced robot geometry. Rather than associating image evidence only with abstract pose-joint parameters, KINO maps image-induced 3D keypoints and forward-kinematics-derived state keypoints into a shared geometric token space, providing a common interface between observation and state. Dense visual evidence and canonical structural cues provide complementary context for state refinement. After each update, KINO regenerates Kinematic Keypoint Tokens from the updated pose-joint estimate through forward kinematics and feeds this geometry back into the decoder. This kinematics-driven feedback supplies the geometry implied by the current state to subsequent refinement, rather than only re-embedding state parameters. Controlled studies demonstrate the benefits of the shared geometric representation and kinematics-driven feedback. On RoboKeyGen and DREAM, KINO outperforms representative baselines, improving real-world AUC by 7.27 and 9.46 points while running at 10.5 FPS (95.53 ms per image).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.