GENBENCH: A CROSS-MODAL MULTI-HOP GE- NOMIC REASONING BENCHMARK WITH CONFIDENCE- COEVOLVING AGENTIC LEARNING
Abstract
Large language models answer genomics questions fluently, but their performance degrades sharply when answering requires chaining multiple pieces of evidence rather than recalling an isolated fact. We introduce GenBench, a 90,379-node, 236,131-edge genomic knowledge graph benchmark built around 210 manually curated seed genes and 5,084 automatically generated, multi-hop, cross-modal question–answer (QA) items, quality-checked by an automated verification pass and audited on a sample by domain experts. On this benchmark, zero-shot Qwen3-8B achieves only 0.2304 accuracy, while GPT-4o and Claude Opus-5 remain below 0.32. Specialized biomedical baselines, including BioReason, Intern-S1-mini, and Evolla, achieve overall accuracies of 0.1676, 0.2807, and 0.2950, respectively, with BioReason dropping to 0.027 on the subset requiring sequence-level evidence. We develop an evidence-aware agent named CoRA (Co-evolving Retrieval Agent) that acquires cross-modal evidence from the graph during reasoning and jointly learns retrieval, trajectory-level credit assignment, and graph confidence through four coupled mechanisms: GRAD (Gated Reward Advantage Distillation), ocbook (open-/closed-book), ROC (Rescored On-policy Credit), and CoCG (Co-evolving Confidence Graph). With supervised warm-start training and graph-based tool access, Qwen3-8B reaches 0.5329 accuracy, increasing to 0.5948 with our full GRPO training procedure, a 158% relative improvement over the zero-shot baseline. The gain comes from evidence access rather than scale, which shows that GenBench resists parametric-only solutions and rewards explicit, confidence-aware retrieval for compositional genomic reasoning. We further identify a substantial multiple-choice/open-ended gap and provide ablations and implementation diagnostics clarifying the training framework's behavior. We release the graph, QA items with complete reasoning chains, and training code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.