Shared Semantic Distinctions and Model-Dependent Expression in Text Encoders
Abstract
Text embeddings support semantic decisions, yet downstream success does not explain which distinctions their vectors encode or how those distinctions are organized. Recent work suggests that independently trained models converge geometrically, raising the question of what meaning they encode in common. We study seven frozen text encoders using controlled sentences that vary six distinctions: success versus failure (outcome), affirmation versus negation (polarity), which event happens first (order), who acts on whom (role), what contains what (containment), and who has which attribute (binding). We compare recovery across constructions and vocabulary, correspondence in responses to semantic changes, and decision-rule transfer as a diagnostic of representational compatibility. Outcome and polarity remain readily recoverable, while relational distinctions show model- and expression-dependent generalization. Binding reveals a clear dissociation: nonlinear readouts improve recovery across all seven encoders on held-out constructions, yet corresponding response geometry coexists with a substantial transfer deficit through the tested linear maps relative to matched within-model readouts. Useful alignment constraints also differ by distinction: matching full responses benefits order, whereas matching their averages performs better for containment. These findings show how apparent semantic deficits can depend on the readout and why geometric correspondence alone is insufficient to establish compatible decision rules. They provide a controlled account of semantic content and organization, clarifying which conclusions about shared meaning persist across linguistic expressions and ways of accessing embeddings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.