acceptodds
Under review as a conference paper at ICLR 2027

Beyond lexical alignment: Testing cross-lingual color representations against human naming

Abstract

Languages partition the continuous colour domain into partly different vocabular- ies, so translation-based alignment alone cannot show whether multilingual rep- resentations capture the colours that speakers associate with colour terms. We introduce a human-referenced cross-lingual probe that maps frozen text represen- tations to language-specific human naming centroids for 1,160 colour terms in 14 languages, collected under a common free-naming condition. The primary eval- uation holds out each of the ten non-English languages with at least 20 reliable terms from both probe fitting and model selection; a sensitivity analysis includ- ing all thirteen non-English languages shows that every model remains above the constant predictor. Five of six models achieve greater relative error reduction than character, token, and frequency controls, and this advantage survives context and exact-string controls in the four models tested. Transfer also persists without English-labelled supervision and across the tested family and script boundaries. The coverage analysis shows where this transfer weakens: standard language holdout leaves the relevant colour regions represented in source supervision, and removing a complete colour region lowers relative error reduction by 0.19–0.25 more than randomly deleting the same number of terms. The results support a transferable mapping from multilingual text representations to human colour ref- erents, while showing that its measured strength depends partly on which regions of the colour domain are represented in source supervision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.