acceptodds
Under review as a conference paper at ICLR 2027

From Cultural Knowledge to Adherence: Diagnosing and Improving LLM Application of Cultural Principles in Realistic Scenarios

Abstract

As large language models (LLMs) increasingly serve users worldwide, they must adapt their responses to diverse cultural principles. Existing benchmarks show that LLMs perform well at identifying explicitly presented cultural principles. However, a troubling discrepancy emerges: recognizing a principle does not necessarily translate into adhering to it in realistic interactions. To investigate the reason under this cultural knowledge-adherence gap and address it, we develop a framework with three connected components. First, we introduce DualCQ, a dual-guided query generation algorithm that jointly optimizes adherence failure under unassisted answering and recoverability under principle-informed answering. Using DualCQ, we construct CulSafe, a benchmark of 10,661 (principle, query) pairs that reveals widespread violations of principles models can correctly identify. Second, guided by cultural intelligence theory, we examine potential bottlenecks underlying these failures. Failures in awareness of cultural sensitivities and contextual principle matching account for a large proportion of remaining unsafe responses. Third, we propose CAPME (Cultural Awareness and Principle Matching Enhancement), which uses a principle-informed teacher to supervise cultural awareness, principle matching, and response generation. Experiments across three model families show that CAPME consistently improves cultural adherence without requiring reference principles at inference time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.