acceptodds
Under review as a conference paper at ICLR 2027

When Language Gains Are Not Semantic: Rethinking Attribution in Language-Supervised Representation Learning

Abstract

Language is increasingly used as auxiliary supervision for representation learning, yet improvements over single-task baselines are often interpreted as evidence that semantic information in language improves the learned representation. This inference is underdetermined by the usual comparison: introducing language supervision simultaneously changes the training objective, optimization dynamics, target geometry, and instance-level supervision. We study this problem as one of experimental attribution and introduce a controlled protocol that separates three distinct questions: whether an auxiliary intervention improves representation learning, whether correct input–target correspondence contributes to the gain, and whether structured targets provide information beyond generic instance-level supervision. Our protocol complements the standard sensor-only baseline with correspondence-destroying derangements and fixed unstructured instance targets, providing explicit controls for competing explanations of an observed improvement. We instantiate the framework in radar representation learning, where language and camera information are available during training but retrieval remains radar-only. On Boreas, correct captions improve Recall@1 by 3.18 points over radar-only training (95% CI [1.34, 5.02]). However, deranged captions and random instance targets also improve performance by 2.75 and 2.17 points, respectively. Correct captions exceed these controls by only +0.43 points (95% CI [−1.66, 2.52]) and +1.00 points ([−0.33, 2.33]), leaving both attribution effects unresolved. The same qualitative attribution result holds for richer camera-VLM descriptions, diagnostic-selected targets, and a reverse-direction replication on MulRan; direct frozen-camera supervision likewise fails to establish stable transfer. We further introduce a pre-training target audit based on cross-revisit stability and between-place collision. Increasing caption entropy from 2.91 to 10.19 bits reduces cross-revisit stability from 0.729 to 0.121, revealing a trade-off between descriptive richness and task-relevant invariance. Our results expose a broader methodological distinction in representation learning: a baseline improvement establishes that a training intervention helped, but does not identify which information within that intervention caused the gain. Semantic attribution therefore requires controls that separately test correspondence and target structure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.