TRACE: Study-Level Transcriptomic Hypothesis Generation with Biological Reasoning
Abstract
Grounded biological reasoning on transcriptomic studies remains a major challenge for genomic and large language models. While existing transcriptomic foundation models learn cell- and gene-level representations, they operate primarily on individual cells and struggle to generate mechanistic explanations for study-level phenomena. General-purpose large language models (LLMs) can generate plausible biological narratives, yet they frequently miss the regulatory pathways and molecular cascades driving transcriptomic signatures. We formulate a task where an LLM is given a disease context alongside descriptions of two biological conditions and must predict the intermediate mechanistic cascades separating them. Using curated condition contrasts from RummaGEO, we construct a benchmark and training dataset of 15,251 human and mouse mechanistic reasoning samples. Each sample maps biological conditions to grounded multi-scale cascades of kinases, transcription factors, and biological processes derived from empirical differential expression enrichment. We then train a TRACE (Transcriptomic Reasoning & Analysis of Condition Expression), an open-source 4B-parameter reasoning model based on Qwen3.5, and conduct rubric-based evaluation of hypothesis generation assessing reasoning trace coherence, final mechanistic summaries, and recovery of top enriched regulatory entities. Across held-out GEO studies and unseen disease contexts, TRACE recovers more empirically supported regulatory terms and mechanistic pathways than open-source and frontier model baselines. Case studies confirm that our model generates supported mechanistic cascades in its reasoning. We release the 15k-sample reasoning dataset, the TRACE model, and the enrichment-to-trace synthesis pipeline to facilitate scalable mechanistic reasoning in transcriptomics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.