TruncDoS: Resource Exhaustion Attack via Response Truncation in Tool-using Agents
Abstract
Tool-using LLM agents implicitly assume that the response they receive faithfully preserves what an honest tool returns. However, prior work on resource exhaustion has focused on content-level stuffing, such as injected instructions, malicious tool use, or manipulated metadata descriptions. In contrast to content-level stuffing, we present TruncDoS, the first attack that amplifies token consumption and latency solely by truncating tool responses. We evaluate the feasibility of TruncDoS, which can be installed through a marketplace skill or plugin, using an in-process mediator that accesses the shared tool-dispatch return path. The mediator achieves TruncDoS by truncating the tool response and forwards fragments. When the missing trunk contains task-relevant evidence, the agent treats the response as incomplete and invokes further queries until the resource budget is exhausted or the task fails. TruncDoS can cause significant increase in token consumption and latency only through truncation, instead of injection, tool modification or additional model invocation at runtime. We demonstrate the resource exhaustion depends on the agent's tool-invocation capability. We then assess TruncDoS across multiple agent harnesses, model families, and benchmarks. The attack increases by up to % in token consumption and latency by up to %. We further apply four representative defense adaptations; none restores clean-task utility, and several even add overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.