MedHarness: A Unified Budgeted Runtime for Reliable Medical Imaging Agents
Abstract
Medical imaging agents still lack a shared runtime reliability layer: under a limited verification budget, deployment must decide when to stop and accept, when to open additional evidence by failure mode, when to repair, and when to reject. Tool-using workflows often freeze a fixed evidence path or a private always-on stack per dataset; those pipelines are costly, still leave many wrong accepts, and cannot share one controller across heterogeneous imaging modalities. We present MedHarness, a registered multi-surface runtime reliability layer that attaches after a proposer's initial prediction: a shared control policy with surface recipe adapters under one packet-route-decision-trace contract, so a new imaging surface attaches by recipe registration to the same stop / open-evidence / repair / reject controller. We evaluate on four registered surfaces (4848 cases) under two complementary tasks: live-pool decision quality and failure-enriched stress efficiency. On a 512-case live pool, MedHarness cuts wrong accepts versus always-on by about 47% and raises accepted accuracy (83.2% vs. 72.9%). On a failure-enriched four-surface stress ranking, the same controller cuts estimated API calls/case under matched context-load accounting by about 71% versus always-on. A held-out three-surface check and an alternate LLM transfer keep the same control ordering. The code is available at https://anonymous.4open.science/r/MedHarness-paper-release-A7F3/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.