acceptodds
Under review as a conference paper at ICLR 2027

When LLM Advice Depends on Values Users Never Stated: Measuring Undisclosed Preference Dependence

Abstract

Large language model (LLM) assistants increasingly give practical advice on decisions whose resolution can depend on how the user prioritizes competing values. Existing evaluations study how models acquire and use preferences; we study what happens before a recommendation-relevant priority has been supplied. We introduce undisclosed preference dependence (UPD): an offline behavioral audit that first tests whether opposing legitimate priorities reverse a model's recommendations consistently across paraphrases and repeated draws, then evaluates separate priority-withheld answers for one-sided commitment without disclosure of that dependency. Across 60 controlled decision families and four model families, the sensitivity criterion held in 188/240 family–model cells (188/216 among estimable cells). Within sensitive cells, human annotation identified UPD in 266/937 eligible substantive withheld answers (28.4%; family-clustered 95% CI 19.9–37.7%). Targeted priority changes produced substantially more movement than the off-axis additions we tested, and the pattern transferred to screened natural-source requests. The assay operationally distinguishes using a supplied priority from handling advice when a recommendation-relevant priority remains unresolved, and provides a scalable, rerunnable offline audit: supplied or constructed scenario–priority contrasts can be reapplied across models and checkpoints, with withheld-answer endpoint judges validated against the human reference, to track what their advice depends on and whether those dependencies are surfaced.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.