acceptodds
Under review as a conference paper at ICLR 2027

ShopEye: Bridging Agentic Data Construction and Active Visual Verification for Fine-Grained E-Commerce AIGC Detection

Abstract

The rapid advancement of AIGC technologies provides a convenient and effective means for product display, driving an exponential growth of AI-generated content on e-commerce platforms. While substantially boosting page views and transaction conversion rates, it introduces uncontrollable generative artifacts, such as distorted product features, severely impairing the user experience. However, existing benchmarks are predominantly confined to discerning whether an image is AI-generated, rather than conducting fine-grained quality assessment on such generated content. To bridge this gap, we present ShopEye, a full-lifecycle agentic framework for fine-grained AIGC defect diagnosis in e-commerce. ShopEye integrates two core components: ShopEye-Agent for automated data production, and ShopEye-Trainer for agentic model training. Specifically, ShopEye-Agent employs an automated multi-agent pipeline to generate and curate diagnostic data. This produces ShopEye-Bench, the first benchmark tailored for low-quality e-commerce AIGC, comprising 5K training and 1K test samples. It covers prevalent defect types in e-commerce scenarios, including product inconsistency, human deformity, unrealistic scenes, and text distortion. Nonetheless, detecting these defects typically requires fine-grained visual understanding. To enhance the model's diagnostic capability, ShopEye-Trainer incorporates diverse visual tools for perception augmentation, and optimizes the diagnostic policy via reinforcement learning. Specifically, it couples Local-Global Dual-View Supervised Finetuning to model fine-grained visual details, and Evidence-Weighted Policy Optimization to incentivize active visual tool invocation for evidence gathering. By combining ShopEye-Bench with our training paradigm, we develop ShopEye-4B and ShopEye-8B, two specialized multimodal agentic models at different scales. Extensive experiments demonstrate that both models consistently outperform strong baselines, validating the effectiveness and scalability of our training framework with only a modest amount of task-specific training data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.