acceptodds
Under review as a conference paper at ICLR 2027

Beyond Single Videos: Benchmarking and Active Evidence Seeking for E-Commerce Cross-Video Reasoning

Abstract

E-commerce videos are information-dense and frequently compared in practice, as consumers compare products and merchants evaluate marketing tactics, making cross-video analysis highly valuable for real-world business analytics. Yet, existing multimodal models primarily focus on single video understanding, with limited ability to perform the cross-video comparisons needed for e-commerce video analysis. To bridge this gap, we introduce **AdsCVR**, the first e-commerce cross-video reasoning benchmark. After filtering, AdsCVR contains 2,483 videos and 6,110 QA pairs across six reasoning dimensions. Extending reasoning across multiple videos fundamentally magnifies the challenge of pinpointing fine-grained evidence amid massive redundant frames, especially in e-commerce videos where visual details, speech, and on-screen text are combined into densely packed content. We propose **AdSeek**, an agentic framework that dynamically allocates visual and audio tools during multi-turn exploration, replacing static uniform sampling with active evidence grounding. While reinforcement learning (RL) can optimize sequential tool use, sparse outcome rewards struggle with long-horizon credit assignment, providing limited guidance on where a trajectory goes wrong or how to correct it. To address this, we propose an **offline trajectory rectification mechanism** that identifies and corrects reasoning errors and missing multimodal evidence in RL-generated trajectories. It then uses them for SFT to strengthen the model’s evidence gathering and reasoning abilities, thereby mitigating biases learned during RL. This mechanism anchors our **rectified bootstrapping pipeline**: an initial RL phase exposes reasoning bottlenecks, the subsequent SFT phase uses rectified data to calibrate model's thinking and action biases, and a final RL phase achieves optimal convergence. AdSeek achieves 74.30% accuracy on the AdsCVR test split, improving upon its backbone (Qwen3-VL-8B-Instruct) by 27.90%. It also generalizes well to the open-domain CrossVid benchmark, validating its effectiveness in active evidence gathering.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.