HR-ViP: Leveraging High-resolution Vision Transformers for Whole-Body Human Pose Estimation
Abstract
Whole-body human pose estimation requires accurate localization of dense body, face, hand, and foot landmarks while maintaining efficiency as the number of people increases. Existing approaches face a fundamental trade-off. Top-down methods apply high-resolution Vision Transformers to individual person crops, providing detailed spatial representations but incurring computation for each detected person. In contrast, shared-feature multi-person methods process all instances within a common image representation but struggle to maintain both accuracy and speed for dense whole-body landmarks. We present HR-ViP, a top-down framework that efficiently leverages high-resolution Vision Transformers for fine-grained whole-body localization and multi-person inference. HR-ViP introduces Train-Time Random Token Subsampling (TRTS), an efficient strategy for adapting ViTs to high-resolution inputs, together with feature-map distillation to transfer their spatial representations to smaller student models. Combined with batched person-crop processing, the distilled models enable efficient multi-person whole-body pose estimation while retaining the benefits of high-resolution representations. In the single-person setting on COCO-WholeBody, HR-ViP achieves 72.1 WB AP at a resolution of . At , an 86M-parameter DINOv3-B student matches the 66.8 WB AP of a 300M-parameter DINOv3-L model while increasing throughput from 91 to 136 FPS. In the multi-person setting, HR-ViP achieves 68.1 WB AP at 50 FPS with ten detected people, substantially outperforming the full-image methods in whole-body accuracy while retaining real-time throughput. These results demonstrate that HR-ViP achieves a strong accuracy–throughput trade-off for both fine-grained whole-body localization and multi-person inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.