Evidence Drift: Evaluating Conditional Reliability in Generative Affective Dialogue Understanding
Abstract
In generative affective dialogue understanding, a correct emotion label can accompany an incorrect sentiment label, cause judgment, or support strategy. Marginal field scores and schema checks do not isolate this conditional failure. We call it evidence drift and quantify it with the conditional evidence drift rate (CEDR). We study paired prediction on MELD, RECCON, and ESConv using Qwen2.5-7B-Instruct. Under supervised joint generation, CEDR is 3.1%, 13.3%, and 36.2% for sentiment, cause judgments, and support strategies, respectively. On ESConv, this risk coexists with a strategy macro-F1 only 0.45 percentage points below that of an evidence-only expert. ESConv CEDR is 73.4% under label-constrained zero-shot generation. The failure is also observed with Mistral-7B-Instruct-v0.3 on RECCON and ESConv. We introduce Conditional Evidence Verification and Deferral (CEVD), which scores the original evidence given the dialogue and predicted emotion. Using a threshold selected on development data, it withholds evidence with weak support while preserving the primary prediction. On the held-out ESConv test set with Mistral, CEVD retains 89.3% evidence coverage and lowers CEDR among retained predictions from 36.4% to 32.6%. It outperforms random deferral at the same coverage (p = 0.0002). These findings support reporting conditional evidence error alongside marginal field performance and evidence coverage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.