acceptodds
Under review as a conference paper at ICLR 2027

AtomDialBench: From Aggregate Scores to Atomic Task Profiles for Dialogue Understanding

Abstract

As organizations increasingly deploy large language models (LLMs) for customer-service review and meeting analysis, model selection demands operation-level evidence rather than merely an aggregate rank. Existing resources offer valuable evaluations but differ in the operations they cover, the outputs they require, and how those outputs are scored. We introduce **AtomDialBench**, a Chinese-English diagnostic benchmark with **16 atomic tasks**, two composite evaluation units, and **3,459** instances. Each task specifies an operation, output requirements, and a scoring procedure; the construction pipeline combines dialogue selection, schema-guided synthesis, and factual checks. Here, atomic means separately specified and scored rather than an independent latent ability. Across **21 LLMs from seven families**, the top five aggregate scores fall within **0.013**, yet task-level rankings differ: the aggregate leader ranks first or joint first on seven of 16 tasks but last on keyword extraction, where it trails the task leader by **0.111**. A 900-response human study yields high inter-annotator agreement (ICC(2,1) ) and a system–human Spearman correlation of . In a controlled comparison with similarly sized training sets, task-aligned fine-tuning yields a higher aggregate score but uneven operation-level changes. Together, AtomDialBench supports aggregate screening, operation-specific comparison, and response-level diagnosis while complementing existing evaluations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.