AnatoGuard: Anatomy-Guided Pose–Semantic Diffusion for Human Amodal Completion
Abstract
Human amodal instance segmentation aims to recover the complete extent of a person, including body regions hidden by occlusion. This task is challenging because invisible regions cannot be directly observed. We observe that the anatomical structure of the human body allows a detected full-body pose to serve as an explicit spatial prior for locating occluded body parts. However, pose alone provides limited information about appearance, scene context, and the continuation of hidden body regions, and its reliability may decrease under severe occlusion. We introduce AnatoGuard, an anatomy-guided pose-semantic diffusion framework for human amodal completion. The detected pose is rendered as a spatial condition and jointly provided with the occluded image and visible-person mask. To complement unreliable pose predictions, a Qwen-based module extracts typed semantic attributes related to appearance, contextual relations, and hidden-body continuation. These attributes are localized using pose-derived body-part supports and transformed into spatial semantic residuals, providing complementary evidence beyond pose. The proposed framework therefore combines anatomical structure with semantic cues for robust amodal completion. We evaluate pose-only and pose-semantic variants through controlled comparisons and component ablations, with particular emphasis on severe occlusion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.