acceptodds
Under review as a conference paper at ICLR 2027

EviVQA: A Dataset and Benchmark for Evidence-Grounded Evaluation of E-Commerce Visual Question Answering

Abstract

E-commerce customer service requires multimodal models to decide not only what to answer, but whether to answer, ask for clarification, abstain, or escalate. Existing benchmarks centered on answer correctness or overall response quality do not fully capture these decisions when evidence is incomplete, conflicting, or beyond available service capabilities. We introduce EviVQA, a Chinese multimodal dataset and benchmark comprising 103,949 instances derived from production e-commerce records. Its retrieval-native inputs combine product images, OCR text, structured attributes, dialogue context, and historical question–answer pairs, while preserving missing information, variant ambiguity, and conflicting evidence. A four-layer annotation scheme connects service scope, evidence sufficiency and information deficits, grounded facts, response constraints, and target actions. These annotations are compiled into decision and evidence certificates, enabling separate evaluation of action selection, factual compliance, and model-judged response quality. On human-annotated, action-balanced examples, we find that higher response-quality scores do not consistently correspond to better action selection, revealing a gap between fluent responses and appropriate customer-service behavior. Combining data curation with certificate-based rewards improves action selection over strong general-purpose and same-backbone baselines, including a 7.6% relative gain in macro-F1 over the strongest fully evaluated general-purpose model. The results further show that improvements in action selection do not necessarily translate into uniform gains in factual compliance or response quality. EviVQA provides a benchmark for studying evidence-grounded customer-service decision making under incomplete information and limited service capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.