Success Is Not Evidence: Verifiable Autoresearch for Cryo-Electron Tomography
Abstract
Autonomous scientific agents are graded by execution success, but execution success is not scientific validity. In cryo-electron tomography (cryo-ET), a picker emits coordinates whether or not particles lie at them, a segmentation returns a mask that can be exactly empty, and a refinement reports a resolution its own half-maps may not support—all with exit code zero. We introduce CRYOFORGE, an autonomous cryo-ET research loop organized around executable evidence: measurements an independent verifier recomputes from stored artifacts, never from tool self-report. A masked admissibility policy acts on the artifact store, and a gated reward scores each trajectory over four evidence tiers with an explicit availability flag, so “not checkable” is never conflated with “bad”. We prove that with sufficiently separated tier weights this reward orders trajectories exactly as a fixed scientific priority does (validity, non-vacuity, localization, resolution), while an ungated scalar can rank the worse policy higher. To our knowledge, CRYOFORGE is the first autoresearch agent for structural cryo-ET: the whole multi-stage subtomogram-averaging reconstruction is the policy’s action space, and the policy is trained on the verifier’s evidence. Experiments carry a published twelve-stage protocol to completion, confirm resolution claims by independent recomputation, and validate picking at particle level (precision 0.57, recall 0.71). The same pipeline generates the controls that test our claims; four weaken or vanish under them.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.