Grounded Product Understanding in Livestream Videos
Abstract
E-commerce livestreams have emerged as an important channel for presenting products to online consumers, often featuring multiple products with relevant information distributed across different moments. This poses significant challenges for downstream product understanding applications, such as product-centric livestream clipping, where models need to identify the product and its relevant segments for information gathering. However, existing benchmarks for general product understanding typically evaluate product retrieval and temporal localization in isolation, leaving the critical correspondence between product identity and temporal evidence largely unassessed. To address this limitation, we introduce GPUB, a large-scale benchmark comprising 3,000 real-world e-commerce livestream instances with quality-controlled multi-moment temporal annotations and a catalog of over 31K fashion products. GPUB supports three evaluation tasks: given a livestream video and a candidate product set, the main task Grounded Product Understanding (GPrU) requires jointly identifying the product being presented and localizing its supporting moments; Product Retrieval and Product Moment Localization serve as two complementary subtasks. Evaluation of existing multimodal models shows that GPrU remains highly challenging, with the best-performing baseline achieving only 10.13 Pair [email protected]. To narrow the performance gap, we further develop UniPro, a unified product understanding model that derives product-aligned and temporally structured representations from shared multimodal encoding, improving Pair [email protected] to 24.58 while achieving 38.81 Joint R@[email protected] on GPrU.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.