acceptodds
Under review as a conference paper at ICLR 2027

GroupRef: Towards Identity-Faithful Video Generation with Multi-Person References

Abstract

Multi-person reference-to-video generation requires preserving every specified identity and assigning each action to the intended individual. Existing benchmarks provide limited coverage of larger groups and often allow appearance descriptions in prompts, making it unclear whether models rely on visual references or textual appearance cues to preserve the specified identities. To address these limitations, we introduce GroupRefBench for groups of three to ten people, with prompts that assign actions to individuals by image index and omit appearance descriptions. Using the reference-to-person correspondences established during identity evaluation, the benchmark further assesses whether each individual performs the assigned actions. To support training with indexed instructions, we construct GroupRef100K, which pairs videos with indexed action annotations and diversified reference images. We further propose Cross-Modal Identity Anchoring to strengthen index-to-reference binding by applying shared learnable anchor embeddings to image-index markers and their corresponding reference-image tokens. Controlled attention analysis provides evidence that the model uses these anchor correspondences for reference selection. The resulting GroupRef models demonstrate promising performance on both GroupRefBench and OpenS2V-Eval.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.