StableCrowd: Spatio-Temporally Consistent Feed-Forward Crowd Reconstruction with Multi-Joint Height
Abstract
Reconstructing crowd motion from monocular large-scene video requires efficiency and spatio-temporal consistency. Among feed-forward methods, per-person video methods recover each person's motion but not the relative positions of the people, while large-scene crowd methods anchor each person in the shared scene through a point regressed as a pixel distance in a crop, a target that varies with the crop and so jitters over time. On the other hand, existing methods that achieve spatio-temporal consistency rely on expensive per-sequence optimization. We present StableCrowd, a feed-forward method that reconstructs crowd motion from such videos. We introduce a new target, Multi-Joint Height, the metric height of every joint above the scene surface. It is crop-invariant, stays stable over time, and can be lifted in closed form to a metric 3D estimate of the joint. We propose Global Trajectory Fusion, which regresses the trajectory in the same space as these estimates, committing directly to absolute positions in the shared scene and covering a sequence of any length. We also propose Canonical Camera Normalization, which absorbs the position and size of each crop into a camera rotation and a zoom, so that close-range single-person data trains the model for large-scene video. StableCrowd is the first feed-forward method to achieve spatio-temporal consistency from large-scene video, zero-shot on VirtualCrowd and WorldPose, showing that per-sequence optimization is no longer required to reach it. Code and models will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.