Seeing Culture When Words Fall Short: Uncertainty-Gated Visual Grounding of Cultural Cues for LLMs
Abstract
Large Language Models (LLMs) have achieved strong performance across a wide range of tasks, yet they remain vulnerable to cultural bias due to the uneven distribution of cultural knowledge in training corpora. This limitation often leads to inaccurate understanding and generation in cross-cultural contexts, particularly when textual cues alone are insufficient. To address this issue at inference time, we propose V-CUE (Visual-Cultural Understanding Enhancement), a visual-space-assisted framework that improves culturally grounded reasoning by incorporating generated visual cues. Specifically, V-CUE first employs an uncertainty detection mechanism to identify potentially unreliable responses. Based on this signal, a text-to-image model generates culturally relevant visual content, including symbols, clothing, and contextual elements. These visual signals are incorporated into the generation process to refine model predictions and enhance cultural grounding. Evaluated on understanding task (CulturalBench) and generation task (CARE) across LLM-text, LLM-reasoning, and LLM-VL systems, V-CUE yields consistent accuracy gains of up to +19 points, while uncertainty gating reduces token cost by 46% with negligible accuracy loss. Further analyses confirm that gains scale with visual–query alignment quality and that V-CUE remains complementary to retrieval-augmented approaches. Without requiring additional training data, V-CUE offers a practical and generalizable multimodal pathway for enhancing the cultural robustness of LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.