Verifier-Guided Skill Evolution for Autonomous Spatial Omics Analysis
Abstract
Spatial omics platforms now pair histopathology images with spatially resolved transcriptomes at scale, yet turning such data into biological conclusions still depends on expert judgment at nearly every step. Recent large language model agents can invoke bioinformatics tools from natural language, but treat each dataset as an isolated episode and discard whatever they learn. We argue the missing ingredient is not a better planner but a usable reward: without a signal separating a biologically sound output from a merely runnable one, an agent has nothing to learn from. We present BioSkill-Evolver, an agent that stores each analysis capability as an executable skill annotated with explicit preconditions and postconditions, moving composability checking from the language model into a symbolic test, and scores every execution with a multi-granularity verifier combining statistical, marker, pathway and spatial checks with a language-model judge. Because this verifier turns an unlabeled analysis into a scalar, it can drive an evolution loop that rewrites failing skills from their own error traces and splits high-variance skills into platform-specific variants under a conservative acceptance rule. Across end-to-end episodes on public Visium, Xenium and Stereo-seq sections, BioSkill-Evolver raises autonomous task success from to and the verification score from to relative to a stateless agent with an identical backbone, while the skill library grows from to skills through three platform specializations. Ablations localize the effect to the evolution loop and the verifier rather than to perception or planning, and the ranking is preserved for three of four alternative backbones. We report this as a single-run, local-benchmark study and state precisely which design mechanisms it exercised.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.