acceptodds
Under review as a conference paper at ICLR 2027

ReMix: Online Evaluation of Evolving Medical AI Systems

Abstract

Medical AI systems built on large language models can change faster than evidence about their performance accumulates. How can we evaluate a new version when only a few patient outcomes are available? We introduce ReMix, an adaptive mixture of historical-borrowing and direct benchmark-score estimators. At each update, paired benchmark evaluations calibrate borrowing in replay-informed estimators, while the new version's benchmark scores provide direct estimates. ReMix adjusts their weights as patient outcomes arrive. We bound its cumulative squared prediction loss relative to the best constituent estimator for any reward sequence. In offline experiments using real patient questions from MedHELM, ReMix reduces mean squared estimation error by 39% after just five post-update patients and cumulative squared estimation error by 29%, compared with evaluating each version from scratch. It also achieves lower average cumulative error than methods for borrowing historical information in clinical trials on both testbeds.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.