Compliance Is Not Discernment: Diagnosing Selective Advice Use in LLMs under Social Framing
Abstract
Effective human–AI and multi-agent collaboration depends on LLMs adopting valid advice while resisting invalid advice. Aggregate following rates cannot distinguish willingness to follow from selectivity for valid advice. We introduce a SuperGPQA-based diagnostic framework to examine selective advice use and test whether cue combinations have effects beyond individual cues. The benchmark pairs valid and invalid advice on stable baseline-wrong model–question pairs in separate calls and tracks harmful regression on baseline-correct pairs. A factorial design crosses source, role, and wording while holding advice content fixed. A probit signal detection model separates the adoption criterion (willingness to follow) from behavioral validity discrimination (selectivity for valid advice). Across eight LLMs, mandatory wording mainly lowered the adoption criterion; supervisor labels amplified this shift beyond additive effects and narrowed validity discrimination under the equal-variance model. Interventions mainly shifted the adoption criterion. Empowered-persona prompting reduced invalid following at the cost of beneficial corrections, whereas native reasoning had framing-dependent effects and sometimes increased harmful regression under mandatory human-source advice. Evaluations of collaborative LLMs should therefore test cue combinations and distinguish selective adoption from general compliance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.