Delegated-Authority Underuse in Large Language Models
Abstract
Imagine an LLM-based digital human helping organize a community event. The human lead tells it, “You have the final say on tonight's schedule.” The plan is ready on screen. Does the digital human confirm it, or send it back for human approval? We find it sends it back. On matched triplets that hold the task, current request, and response options fixed and differ only in the dialogue assigning responsibility, the models answer correctly 99.0% of the time when the human holds final approval, but only 33.3% when final authority is delegated to the model. In 96.25% of model-higher errors, the models hand the decision back to the human. We term this one-sided failure delegated-authority underuse (DAU). Prior work on deference and sycophancy has not asked whether models retain a final right explicitly given to them. We address this open question with the Delegated Authority Response Controller (DARC), a lightweight inference-time method that makes the authority direction explicit and binds it to a response policy before generation. Across five English models, DARC raises average forced-choice accuracy from 76.0% to 99.3% and model-higher accuracy from 33.3% to 97.8%. The same repair holds in French and Chinese and in human-evaluated free-form responses, and this consistency shows that recognizing authority and exercising it are distinct capabilities. Delegation gives models the final word, yet they return it; making the assignment explicit is enough to make them keep it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.