acceptodds
Under review as a conference paper at ICLR 2027

Contextualizing Biological Language Models across Modalities via Logit-Space Contrastive Alignment

Abstract

Pretrained biological language models encode rich probabilistic structure through native token distributions, but multimodal adaptation typically aligns hidden representations or replaces the native likelihood interface with task-specific heads. We introduce LogiCA (Logit-space Contrastive Alignment), whose central methodological change is to align the context-conditioned native token distributions themselves rather than pooled latent representations. Using native-head-preserving cross-modal adapters, LogiCA couples models with distinct vocabularies and tokenization schemes without requiring a shared tokenizer, decoder, or embedding space. Contrastive learning is performed directly through each model's original token-likelihood interface, retaining token-level scoring, interpretation, and generation. Across protein–ligand, TCR–peptide, and drug-resistance tasks, LogiCA outperforms matched latent-contrastive and conditional-MLM baselines, particularly in low-homology, mutation-local, and data-scarce regimes. These results support native-logit alignment as a general strategy for making pretrained biological token distributions context-sensitive without sacrificing their probabilistic semantics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.