BioStitch: Gated Cross-Attention for Biological-Language Modeling
Abstract
Integrating biological foundation models with large language models (LLMs) is crucial for a broad range of biological tasks, from sequence understanding to text-guided biological generation. However, existing approaches typically connect these models through input-space fusion, without fully exploiting the rich representations distributed across their internal layers. To understand how pretrained representations can be better connected beyond input-space fusion, we systematically probe the layer-wise interaction between biological models and LLMs. We find that, when transferring DNA representations into LLMs, intermediate layers provide more effective features on average than conventionally used final-layer representations. Guided by this finding, we propose BioStitch, a versatile framework that stitches biological foundation models and LLMs through multi-level gated cross-attention, allowing hidden states in one model to selectively query internal representations from the other. On biological-sequence-to-text tasks, BioStitch delivers strong improvements on DNA variant-effect benchmarks, protein question answering, and zero-shot single-cell annotation, without multimodal pre-training. On text-conditioned protein generation, BioStitch substantially improves text-sequence alignment over conventional input-space conditioning. BioStitch is also highly efficient in the measured long-DNA setting, reducing peak training memory by up to and LLM prompt KV-cache usage by . All these results indicate that BioStitch is an effective and efficient framework for connecting specialized biological and language representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.