acceptodds
Under review as a conference paper at ICLR 2027

ComplaintBench: A Dataset and Benchmark for Evidence-Grounded Payment Complaint Assessment

Abstract

We introduce ComplaintBench, a dataset and diagnostic benchmark for evidence-grounded risk assessment in Chinese-language payment complaints: 87,359 records with 286,582 screenshot references, each carrying a narrative, chat and payment screenshots, a binary positive / non-positive decision, and one of 32 operational categories. The decision policy places ordinary disputes and insufficient-evidence cases in the non-positive class, so “unverifiable” is a legitimate outcome rather than an abstention. A two-stage reference pipeline extracts source-linked evidence from text and screenshots with a frozen vision-language model and then classifies from the resulting text view; we compare it with six hosted services under matched evidence and under direct raw input. At 16.93% positive prevalence, accuracy is not a usable criterion—a constant non-positive control scores 83.1% accuracy, above every evaluated system, with zero recall—so we report BalancedSuccess alongside recall and false-positive rate. Fine-tuning determines the error direction: the complaint-fine-tuned model reaches 75.7% recall at a 17.0% false-positive rate, against 74.8% recall at 43.5% for the same model without complaint-specific fine-tuning (BalancedSuccess 79.4 versus 65.7), whereas every hosted service buys 71–100% recall only at a 34–65% false-positive rate that concentrates on insufficient-evidence records. A de-identified 20,000-record subset with its extracted evidence is released as supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.