Understanding and Evaluating Cultural Norm Grounded Reasoning in LLMs
Abstract
Cultural intelligence in large language models (LLMs) requires not only knowing cultural norms but also reasoning with them under context. Yet existing benchmarks predominantly evaluate cultural knowledge recall, leaving open whether models can effectively invoke acquired norms in multi-constraint reasoning scenarios. We introduce CultureForest, a benchmark for Cultural Norm Grounded Reasoning that explicitly tests this ability. Each of its 5,378 examples is grounded in atomic cultural norms, enabling verifiable and attributable evaluation across 8 domains and 53 countries/regions. CultureForest supports progressive evaluation from multiple-choice (Easy) through binary judgment (Medium) to open-ended generation (Hard), accompanied by a lightweight verifier for scalable open-ended assessment. Extensive experiments reveal that even top-tier models degrade substantially in open-ended settings, with pronounced cross-region disparities. Through multi-faceted analysis, we uncover consistent patterns: (1) test-time reasoning yields limited gains and may exacerbate inequity; (2) models share convergent regional preference structures; (3) responses grow markedly conservative under stricter norms; and crucially, (4) by disentangling cultural knowledge acquisition from reasoning, we demonstrate that the primary bottleneck is not knowledge availability but its effective utilization during reasoning. These findings point to a necessary shift from knowledge-centric evaluation toward measuring knowledge-grounded reasoning in cultural scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.