acceptodds
Under review as a conference paper at ICLR 2027

Exploring How LLM-Guided Scientific Discovery Can Find Solutions to Long-Standing Challenges in Disease Risk Prediction from Genetic Data

Abstract

Predicting a person’s risk of disease from their genome works best for people whose genetic ancestry is well represented in the data used to build the model, and worst for everyone else. Because data from underrepresented ancestry groups are scarce, genetic risk models in precision medicine routinely borrow data from other, better-represented groups. How much to borrow, and from whom, is decided by data-borrowing schemes that are designed by hand and apply one global rule to every genetic variant and every person. We ask whether LLM-guided scientific discovery can find a better scheme. An LLM writes candidate schemes as short Python functions that decide, for each person and each variant, how far through continuous genetic ancestry space to borrow labeled subjects. A fixed evaluator scores each candidate on real prostate cancer data, and the best-scoring scheme is locked before any validation data are touched. With East Asian ancestry as the underrepresented (focal) group and discovery performed on one cohort, the locked scheme achieved the highest held-out discrimination (AUC) in three independent validation settings drawn from three different data sources, compared with four established borrowing schemes (Mixture Learning, Independent Learning, and two transfer-learning schemes). Pooling all 112 held-out focal subjects, one-sided paired tests favored the discovered scheme over every comparator (p < 0:05 for all four). An ablation shows that weakening the evaluator during discovery degrades the held-out performance of the resulting scheme, indicating that the evaluator, not the LLM alone, drives what is found.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.