BioBLIP: A Multimodal Architecture for Zero-Shot Task Transfer in Genomic Variant Interpretation
Abstract
Developing scientific hypotheses in biology requires integrating heterogeneous evidence from DNA sequence, gene context, protein function, and prior literature. Existing multimodal AI systems expose biological evidence to LLMs through textification or by projecting biological embeddings into fine-tuned language models. However, these models' performance is limited to the specific tasks for which they are fine-tuned. Here we present BioBLIP, a multimodal Q-former based architecture which leverages biological embeddings and a LLM to generalize to unseen genomic tasks without fine-tuning. The key to BioBLIP is a new model architecture that integrates four data modalities – DNA, genes, proteins, and text – through a master Qformer model, which consolidates the modality-specific information into multimodal tokens for a frozen LLM. BioBLIP is pretrained on the task of human single nucleotide variant annotation (85.2% overall accuracy). Without fine tuning and only 66M trained parameters, pretrained BioBLIP achieves 70.2% Top-1 accuracy in variant prioritization, as opposed to 64.9% accuracy for LLMs given matched biological evidence in text and 59.3% for 40B parameter genomic foundation models. In target gene prediction, BioBLIP outperforms matched-text LLM baselines (60.50% versus 43.70%) while producing rich, transparent traces. Overall, BioBLIP demonstrates strong capabilities to leverage learned multimodal representations for robust zero-shot transfer to multi-step inference tasks in genomics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.