Identity Meets Interaction: Diagnosing the Joint Identity–Interaction Reliability Gap in Reference-Guided Handheld Product Insertion
Abstract
Recent advances in instruction and reference-guided image editing enable models to follow complex instructions, transfer the appearance of reference content, and generate compelling visual results. However, editing guided by a reference object image requires more than a plausible composite: the object's identity-defining details also need to be preserved. This requirement is especially important in e-commerce, where product text, logos, and other identity cues carry commercial value. Existing models and benchmarks primarily focus on instruction following and overall visual quality, leaving the joint satisfaction of product identity and plausible hand-object interaction largely outside quantitative evaluation. To fill this gap, we introduce HandHeldBench, a diagnostic benchmark comprising 1,000 controlled-synthetic samples and 1,000 real-captured samples. We further propose HandHeldEval, an evaluation protocol for product scale, structural and high-frequency detail, identity similarity, and hand-object interaction plausibility. Results across mainstream general-purpose and task-specific editors on HandHeldBench show that current models cannot consistently achieve both plausible hand-object interaction and faithful preservation of product identity details. We define this recurring limitation as the Joint Identity-Interaction Reliability Gap. Beyond benchmarking, we construct HandHeld-55k, a training dataset of 54,923 validated editing quartets, and study reference-guided product-detail restoration as a strong post-processing baseline. Together, these components provide a systematic foundation for analyzing identity preservation and physical interaction in handheld product insertion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.