Asking Forever: Turn Amplification via Clarification-Seeking Directions in LLMs
Abstract
Multi-turn interaction length is a dominant factor in the operational costs of conversational LLMs. In this work, we present a new failure mode: *turn amplification*, in which a model consistently prolongs multi-turn interactions without completing the underlying task. We show that an adversary can exploit clarification-seeking behavior—commonly encouraged in multi-turn conversation settings—to consistently prolong interactions. Moving beyond prompt-level behaviors, we take a mechanistic perspective and identify query-independent (i.e., *universal* across prompts) activation directions associated with clarification-seeking responses. Unlike prior attacks that rely on per-turn prompt optimization, exploiting these directions requires no such optimization, and their effect persists across tasks. To assess whether it constitutes an exploitable vulnerability, we instantiate two well-established threat models—supply-chain attacks via fine-tuning and runtime attacks through hardware-level fault injection—and show that both reliably induce turn amplification across prompts and tasks. Across multiple LLMs and benchmarks, our attack substantially increases turn count while keeping individual responses benign. We evaluate potential countermeasures and show they offer limited protection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.