The illusion of Task Success: Structured Behavioral Diagnosis in Multi-turn LLM Interactions
Abstract
Large language models (LLMs) achieve strong performance on standard benchmarks. However, existing evaluations remain dominated by aggregate metrics that reduce dynamic interactions to coarse scores. In this work, we show how conventional evaluation approaches may create an illusion of task success by masking severe late-stage failures across multi-turn interactions. We introduce RHCA, a structured framework informed by 126 literature sources that tracks model behavior across four core dimensions: reasoning transparency, helpfulness, consistency, and context alignment. To support RHCA at the trajectory level, we also introduce the Behavioral Trajectory Verifier (BTV), which is designed to capture severe failures and performance drops across turns. Our evaluation includes an in-depth music programming study and cross-domain evaluations using tasks created by 16 AI practitioners. While conventional approaches, including automated LLM judges, may miss critical late-stage failures, BTV can flag failures and score drops across turns. These findings highlight the value of structured trajectory evaluation for assessing LLMs in multi-turn interactions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.