acceptodds
Under review as a conference paper at ICLR 2027

Who is a Better Architect: Subjective and Objective Quality Assessment for Talking Human Interaction Pipeline

Abstract

Interactive digital humans are commonly built as cascades of specialized components, known as learnware (LW), offering lower response latency and smaller deployed models relative to omni models. Existing quality assessment, however, focuses primarily on individual learnware and lacks a system-level framework for evaluating the overall interaction. We integrate four learnware central to digital human interaction—a large language model (LLM), text-to-speech (TTS) synthesis, text-to-image (T2I) character generation, and a talking-human (TH) generator, or Talker—into the LT3 pipeline. To study interaction quality from the perspective of learnware, we introduce THQA-LW, a dataset of 10,000 single-turn interaction videos, together with a subjective protocol that separately scores each learnware and the overall interaction experience. We then analyze how learnware quality relates to overall interaction quality under controlled generation conditions. Finally, we propose LT3-FAMT, a five-task quality predictor that combines Qwen2.5-Omni Thinker with complementary multimodal feature extractors. LT3-FAMT achieves state-of-the-art performance across multiple benchmarks, supporting automatic assessment and diagnosis of quality bottlenecks in interactive digital humans.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.