Hypothesis Generation and Semantic Attribution in Inverse Perturbation Modeling
Abstract
Understanding gene regulatory relationships at scale can help identify intervention points and anticipate unintended effects. However, connecting a black-box model's learned computation to biological knowledge remains difficult, particularly for unfamiliar cellular states. We investigate whether a sparse, typed gene network can generate inspectable hypotheses for steering malignant transcriptional programs toward expression patterns associated with favorable prognosis. Using LINCS L1000 CRISPR/Cas9 profiles paired with matched plate controls, we train a graph neural network to recover the perturbed gene from baseline and endpoint expression. Its nodes combine ProTrek-derived protein representations with expression features; its edges encode signed regulatory relationships and other gene interactions. Our focus is how representations learned from perturbation responses use these identities and relationships to support individual nominations. We then query the model in a specified cancer-cell context with endpoints that upregulate favorable prognostic markers and downregulate unfavorable ones. At this fixed request, integrating from mean projected protein representations to actual representations allocates a nomination score difference through node states, individual layer-edge messages, and residual connections. In detailed A549-PYGL and A549-CCDC86 cases, supporting and opposing contributions coexist. Connected supporting structures suggest hypotheses involving glycogen-associated regulatory/metabolic propagation and physical interactions associated with ribosome biogenesis. Preserving gene and relation identities makes these learned computational dependencies explicit, allowing comparison with biological knowledge and framing mechanistic hypotheses for experimental investigation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.