To Ask or Not to Ask: Backdooring the Learned Consultation Policy of Collaborative LLM Agents
Abstract
Small–large model collaboration is an emerging paradigm for agentic systems that pairs a powerful but expensive large model with a smaller, lower-cost model to balance performance and inference cost. The small model executes each task and uses a learned policy to decide whether to consult the large model. We show that this learned consultation policy introduces a new attack surface. Specifically, we present CONFUNDO, the first backdoor attack on this learned consultation policy. By appending a benign natural-language instruction to the input as a trigger, we manipulate the consultation policy in two opposing directions. Denial-of-Wallet (DoW) forces unnecessary consultations, substantially increasing inference costs, whereas Denial-of-Expert (DoE) suppresses necessary consultations, degrading task performance. With only approximately 4% poisoned training data, DoW achieves a 100% consultation rate on triggered inputs in most evaluated settings while preserving behavior comparable to that of a benign model on clean inputs. Conversely, DoE reduces the consultation rate on triggered inputs to 0%, lowering accuracy on GSM8K from 0.800 to 0.595 and increasing the insecure-code generation rate on security-sensitive coding tasks from 0.356 to 0.488. Existing prompt-injection and jailbreak attacks fail to produce comparable manipulation of the consultation policy, and no single operator-side defense we evaluate detects both payloads. Our findings reveal that learned consultation policies can undermine both the economic and performance benefits that small–large model collaboration is designed to provide.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.