PoseStage: Unifying External Pose and Internal Structure Control in Multi-Object Scenes
Abstract
Controllable image generation can place objects, specify their orientation, or articulate a human body, but in scenes with several open-category objects these controls are rarely available at the same time. We present PoseStage, a unified framework that controls both the external pose and the internal structure of multiple objects from two 2D geometric maps. PoseStage supports three novel pose-and-structure tasks within one framework: generation, reference-guided generation, and editing. A model trained only on standard conditional generation handles the latter two at inference time, without task-specific training, by extending its input sequence. To prevent the appearance and text of different instances from mixing, we apply a training-free instance-routing mask at inference that removes direct attention between instances and reduces attribute confusion. Training a single model to control both aspects requires images annotated with external pose and internal structure at the same time, which existing datasets do not provide. We therefore build a hybrid dataset of real-world and synthetic images that provides aligned annotations for external pose and internal structure, and a benchmark that evaluates both controls, on which PoseStage outperforms existing methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.