MuPeT: Multi-Person Transformer for 3D Reconstruction from Broadcast Sports
Abstract
High-level sporting events are often recorded from multiple viewpoints, providing abundant visual data to train models for 3D reconstruction and human pose estimation. However, these settings remain a significant blind spot for current models: they involve highly dynamic scenes with limited texture and repetitive backgrounds, captured by uncalibrated, long-range cameras. In this paper, we present MuPeT (Multi-Person Transformer), a model to reconstruct player poses and impacts in 3D under severely challenging conditions. Since training a Visual Geometry Guided Transformer (VGGT) requires precise 3D supervision, absent in real sports data, we train it exclusively with synthetic renderings. The key enabler is the input modality: instead of pure RGB, which would present a large sim-to-real gap, our model consumes human segmentations and lines, which are easily rendered or computed from real data. We thus leverage the strengths of pretrained models while enabling purely-synthetic training, to learn the geometric biases inherent to these environments. We show that MuPeT successfully reconstructs difficult impact scenarios where previous approaches fail.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.