acceptodds
Under review as a conference paper at ICLR 2027

PAIQ: Patch-Aligned Semantic Injection with an Orthogonal Bridge

Abstract

Language-aligned and self-supervised visual encoders offer complementary strengths in semantic abstraction and spatial detail. Making use of this complementarity requires incorporating language-aligned semantics into local visual representations. We introduce PAIQ, a visual fusion method that enriches DINOv3 patch features with SigLIP features through patch-wise residual updates. PAIQ uses DINOv3 features as the base representation and projects both visual streams into the language-model dimension. Learned content similarities and a finite-step Sinkhorn procedure determine how SigLIP features are aggregated for each patch. A shared orthogonal transformation maps the difference between the aggregated and base features, and the transformed residual is added back with a fixed scale. This design learns update directions while preserving residual norms and pairwise angles, providing an explicit geometric constraint on semantic enrichment. Both visual encoders and the language model remain frozen; only projection and fusion parameters are trained, while the fused sequence retains 196 visual tokens. Across five language backbones and four evaluated datasets, PAIQ achieves higher correctness and lower hallucination severity than single-encoder baselines in most settings. PAIQ also compares favorably with mainstream multi-encoder and token-compression baselines, including an adapted CoME-VL baseline in the 2B setting. These results demonstrate that geometrically constrained residual updates provide an effective way to integrate complementary visual representations while maintaining a compact visual interface.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.