CONTACT: A Human-Grounded Benchmark and Surprisal-Based Predictive Scorer for Conversational Naturalness
Abstract
Spoken dialogue evaluation has advanced from response correctness to turn- taking, emotional responsiveness, and conversational quality. Yet a central chal- lenge remains: determining whether automatic scores reflect the naturalness peo- ple experience across both targeted interaction failures and spontaneous conver- sation. We introduce CONTACT, a human-grounded benchmark and a predictive scorer to address this challenge. The benchmark combines live human–human conversations with elicited turn-taking and affective disruptions, and free-form human–AI dialogues. Human ratings anchor the assessment of turn-taking, af- fective, and overall naturalness across the benchmark’s subsets, connecting diag- nostic control with naturalistic evaluation. Alongside this resource, we propose a surprisal-based scorer trained on approximately 1,600 hours of natural dyadic speech to anticipate future turn-taking and affective states. Surprisal of the ob- served speech given a previous context yields dimension-specific scores and an overall composite without requiring a matched reference conversation. Results suggest that the proposed scorer provides a higher correlation with the perceived naturalness provided by the benchmark than current full-duplex evaluation sys- tems. Together, the benchmark and scorer connect targeted diagnosis, human- perceived interaction quality, and natural-dialogue predictive learning, providing a reusable basis for evaluating conversational naturalness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.