LHC-Net: Bridging LLM Priors and Hypergraph Structure via Channel Retrieval for Biochemical Forecasting
Abstract
Multivariate Time Series Forecasting (MTSF) serves as a core instrument in AI for Science (AI4S) for analyzing complex biochemical processes. However, practical applications face a fundamental dilemma: balancing the physical deficiency caused by Channel Independence (CI) against the noise overfitting induced by Channel Dependence (CD). While CI mitigates high-dimensional noise, it severs essential physical coupling; conversely, CD models global dependencies but introduces spurious correlations, leading to redundancy and overfitting. To address this, we propose LHC-Net (LLM-initialized Hypergraph Channel-Retrieval Network), a novel paradigm integrating scientific domain knowledge with deep learning. Unlike methods employing Large Language Models (LLMs) as direct numerical predictors, we leverage their general scientific knowledge to extract sparse hypergraph structural priors, characterizing high-order mechanistic dependencies among variables. To mitigate the risk of LLM hallucinations, we define this topology as learnable parameter matrices. This allows dynamic correction of priors via gradient backpropagation from observational data, achieving robust integration of mechanistic priors and empirical evidence. Furthermore, we design a temporal channel retrieval strategy that enables the model to access critical variable information from the mechanistic hypergraph on-demand, guided by future temporal contexts. This structured channel fusion strategy restores the necessary physical coupling of dynamical systems while effectively shielding against irrelevant interference. Extensive experiments on a million-scale real-world biopharmaceutical dataset demonstrate that LHC-Net significantly outperforms state-of-the-art models in accuracy. Moreover, it successfully recovers reaction paths consistent with biochemical principles. By grounding LLM priors in observational data, LHC-Net effectively mitigates hallucination risks, enhancing the application potential of AI4S. Code is available at https://anonymous.4open.science/r/LHC-Net-8B41
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.