acceptodds
Under review as a conference paper at ICLR 2027

CogCanvas: A Benchmark for Evaluating Multi-Subject Reference-Based Image Generation

Abstract

Multi-subject reference-based image generation requires jointly preserving multiple human identities, binding per-person objects and fashion items, and respecting a specified background scene, a regime where current diffusion models remain brittle. Existing benchmarks evaluate only one axis at a time and none jointly captures multi-identity composition with human-object interaction, background grounding, and spatial plausibility. We introduce CogCanvas, a benchmark of 1,952 curated reference images spanning 100 celebrity identities, 115 distinctive objects and fashion items, and 29 real-world background scenes including Vietnamese landmarks, from which we construct 1,361 compositional prompts covering 2–5 person group sizes. The curation pipeline combines DINOv2-based deduplication, automated filtering, human review, and structured interaction and position annotations that make individual relations checkable. CogCanvas supports three tasks, reference-based multi-human-object generation (primary), text-to-image compositional generation, and reference retrieval, under a unified six-axis evaluation protocol. We introduce two metrics tailored to the multi-reference setting: BG-Sim, which scores background fidelity on SAM 3-masked regions via DINOv3 feature similarity, and Attr-VQA, which uses a multimodal LLM to verify per-subject attribute binding and inter-person interactions against the structured graphs. Benchmarking five existing methods and a training-free sequential-inpainting reference pipeline shows that object and fashion fidelity drops sharply as group size grows from two to five. A 46-image human study finds moderate agreement between Attr-VQA and human judgments (Spearman ρ = 0.645).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.