acceptodds
Under review as a conference paper at ICLR 2027

Do LLMs Walk the Talk? A Benchmark for Probing Appraisal-Action Inconsistency in Value Systems

Abstract

Existing LLM value benchmarks often rely on universal dimensions that treat values as independent factors, overlooking their hierarchical dependencies, cross-value relations, and context-dependent expression. We introduce a news-grounded, comparative benchmark for evaluating LLM value appraisal, action, and the inconsistency between them across China, the U.S., and the U.K. We propose the Human-AI-Delphi Value Construction Framework (HAD) to construct country-specific normative reference systems that combine a hierarchical upper structure with networked relations at the lower level. Based on these systems, we develop a dual-perspective evaluation method for value appraisal and action. Presupposition Fallacy Questions (PFQ) assess how LLMs respond to value-laden presuppositions, while Scenario-Induced Questions (SIQ) assess how they select and justify actions under situational pressure. Each PFQ–SIQ pair is derived from the same news article and fine-grained value indicator, enabling paired measurement of Appraisal–Action Inconsistency (AAI). The benchmark supports open-ended responses and continual updates. Experiments across LLMs reveal distinct patterns in value appraisal and action. PFQ shows stronger variation across models and country contexts, with an asymmetric localization pattern, whereas SIQ is more stable across models and countries. AAI is widespread across models and countries and varies substantially with both model and context.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.