acceptodds
Under review as a conference paper at ICLR 2027

Pseudo-Deliberation: Measuring and Mitigating Value–Action Gaps in LLM Dialogue

Abstract

Large language models (LLMs) can articulate values that do not reliably translate into their behavior, a discrepancy known as the value–action gap. In this work, we argue that the value–action gap persists even under explicit reasoning, revealing a deeper failure mode we call “Pseudo-Deliberation”: the appearance of value-aware reasoning without corresponding behavioral alignment. We develop a framework for measuring and mitigating this failure across the value–reasoning–action pipeline. First, we introduce VALDI, which evaluates value–action alignment by tracing value commitments across self-surveys, explicit reasoning, and generated dialogue. VALDI instantiates this framework with 4,941 human-centered scenarios across five domains, three elicitation tasks, and five complementary measures of value adherence. Applying VALDI across instruction-tuned and reasoning-native LLMs reveals persistent value–action gaps and enables us to localize where value commitments are lost across the pipeline. Building on this decomposition, we introduce VIVALDI, a multi-agent intervention targeting both reasoning-stage value suppression and reasoning-to-action translation failures. VIVALDI decomposes deliberation into competing value-specific actions via debate and explicitly repairs commitments that would otherwise be lost before dialogue generation. Across models, this structured intervention improves value–action alignment by approximately 19–21% relative to standard chain-of-thought prompting. Together, VALDI and VIVALDI provide a framework for diagnosing and intervening on failures in translating articulated values into model behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.