Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Abstract
Meaning identity—whether two sentences say the same thing after the wording has changed—is treated throughout retrieval and RAG as a geometric fact about independently encoded sentence vectors. We test a different picture: identity is computed when both sentences occupy the same forward pass, not shipped as a cheaply readable property of either sentence’s embedding. On a common 2,000-pair English PAWS test set, linear readouts of four frozen joint passes reach AUC 0.900–0.946 with 1,200 labels, versus 0.622–0.646 for full-width nonlinear readers of independent states. With all 49,258 training pairs, those readers improve to 0.736–0.786, and cross-attention over frozen token sequences reaches 0.868 on a 14B backbone, while the joint readouts remain 0.922–0.960. Across a separate overlap-matched survey of 30 decoder checkpoints, joint exceeds an independently selected product/difference probe in 29 rows and partner shuffle collapses to chance; a pair-disjoint replication reverses the sole GPT-2-large ordering, although its intervals overlap. The same joint-over-independent pattern recurs in bidirectional encoders, encoder–decoders, cross-language and nonce controls, and controlled role, negation, and quantity relations. It is not the generic consequence of scoring jointly: BGE rerankers enter the identity band, while MS-MARCO and Jina rerankers do not. Nor is it an immutable architectural limit: fine-tuning BGE or applying rank-16 LoRA to E5-Mistral can install a strong independent criterion by changing the encoder. The sharper result is route-dependent teachability. A 1.5B joint student and a trainable independent student receive the same 9,000 teacher-scored pairs and update budget, yet reach English PAWS AUC 0.942 and 0.760; six independently trained models also produce strongly concordant, symmetric, and graded joint scores. Across the tested models, readers, and supervision budgets, sentence geometry remains useful for aboutness, but propositional identity behaves as a recurring relation made cheap to read by computation over the pair.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.