acceptodds
Under review as a conference paper at ICLR 2027

The Retrieval–Grounding Tradeoff in Cross-Modal Connectors, and How to Navigate It

Abstract

We ask whether a topology-preserving, invertible flow connector traces a better alignment-vs-structure frontier than projection heads for cross-modal (image<->text) translation, under the Aristotelian Representation Hypothesis (Gröger et al., 2026) that neural representations share local-neighborhood but not global-metric structure. Across ten connectors and two pre-registered testbeds — within-CLIP (pre-aligned) and cross-model (independent encoders) — the flow does not win: transport flows land off the target manifold; a guided conditional-generation flow buys single-shot retrieval but not manifold fidelity; and a one-line InfoNCE term on the same MLP beats every flow (42.9 vs. 27.9) at a fraction of the cost. Our central result is causal: a connector's grounding (how decodably image-specific its output is) is driven by proximity to the target manifold, which we establish with a projection/shuffle/intervention triad — so optimizing a connector for retrieval pushes its outputs off the manifold, a marginal manifold cost of discriminative pressure that is a hard retrieval–grounding exclusion when the spaces are pre-aligned and softens to a frontier otherwise. The mechanism is prescriptive: a train-time nearest-real-target pull on a contrastive objective navigates the tradeoff, filling the previously-empty top-right and Pareto-dominating every learned connector in both regimes (+11 R@1 over regression at equal grounding), and — the key control — beating post-hoc DeCap-style and subspace projections that target the same proximity. We scope all claims and release full pre-registration, code, and per-seed results.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.