acceptodds
Under review as a conference paper at ICLR 2027

Fair Fact-Checking: Closing the Cross-Lingual Gap in LLM Factual Judgement with RoSh

Abstract

Misinformation on social media remains a critical problem, and more and more people settle it by asking a language model instead of a fact checker, about anything from a claim circulating on a platform to a plain question of fact. Whether models judge such claims reliably is still debated; whether they judge them equally well in every language people ask in has gone almost unasked. We test eight models from five families, 3B to 70B, on 1,500 claims that mix fact-checked news with encyclopedic facts and exist in identical form in eight languages. On every model English is judged better than the other languages, and the gap is widest on the smallest models, where Llama-3B on Arabic is no better than guessing. Existing remedies either retrain the model on more multilingual data or fit an unconstrained map between language representations, and neither asks whether the model already holds the answer and simply fails to say it. It largely does: a linear probe recovers the truth from the very activations the model fails to express. We propose RoSh, a per-language shift and rotation of the residual stream, computed in closed form at three layers, with no training and no weight modified. It improves every model and closes 75% of the gap on average, and it helps most where the model was worst: the two smallest models end close to English, Arabic on Llama-3B goes from chance to nearly the English level, and a fifth fewer of the claims answered correctly in English are lost in translation. What remains is no longer a read-out failure: afterwards the head recovers at least as much of what is encoded outside English as it does in English. An unconstrained map fitted on the same pairs falls below the untouched baseline, so the orthogonality constraint is doing the work, and every model clears a scrambled-correspondence control and ten further controls. On the two benchmarks of the closest inference-time method, latent-space intervention, run with its own data and metric code, RoSh's gains are five to thirteen times larger.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.