Structure-Guided Alternating Inference for Monocular Robot State Refinement
Abstract
Estimating the pose and joint angles of an articulated robot from a single RGB image supports spatial reasoning without access to its internal measurements. Existing approaches use learned visual representations to predict the robot state, sometimes with iterative refinement. However, articulation and global pose can produce similar image changes, making it difficult to determine which variable should correct a remaining discrepancy. To address this, we introduce Structure-Guided Alternating Inference (SGAI), a training-free method that formulates state refinement as bounded geometric optimization using robot geometry and perceptual cues extracted by visual foundation models. Specifically, SGAI solves separate articulation and pose subproblems while holding the other block fixed, preventing compensating changes between the two blocks within each solve. Point-to-link, silhouette, and projection anchor constraints assign geometric evidence to the appropriate updates and limit corrections driven by imperfect depth and foreground predictions. Re-rendering and renewed geometric associations allow each block to use the shape or placement corrected by the other. Experiments on DREAM show improved 3D localization in most tested settings across two initializers and three robot models, using a shared configuration without additional training or initializer- or robot-specific retuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.