acceptodds
Under review as a conference paper at ICLR 2027

When Can Expert Corrections Be Trusted? Evidence, Limits, and Risk-Controlled Decisions Across Environments

Abstract

When a strong reference model and an automated expert disagree, the same correction rule can help on one patient's slide and harm on another. We ask when such corrections can be trusted across environments, and propose a unified framework separating the construction of corrective evidence from the decision to act on it. On the evidence side, we introduce CARE-Cell, a cell-classification architecture combining frozen tissue representations, local nuclear features, and target-conditioned context to build both a strong reference predictor and complementary corrective candidates. On the decision side, a risk-controlled layer determines when to revise, retain, or defer. Theory explains why both parts are necessary: we construct observationally indistinguishable environments whose optimal correction decisions are opposite, yielding a worst-case regret lower bound for any router measurable with respect to reference outputs, expert outputs, and confidence. The decision layer accordingly offers two complementary guarantees. With target-environment labels, an independent randomized audit certifies positive aggregate utility on the remaining queries with a finite-sample error guarantee, retaining the reference predictions when certification fails. Without such labels, labeled calibration environments set an accept/defer threshold that, under exchangeability, keeps a new environment's retained-error fraction below a prescribed limit with a marginal probability guarantee. We evaluate on a consecutive dermatopathology cohort collected without case selection based on image appearance or model performance, together with Lizard, PanNuke, Camelyon17-WILDS, and iWildCam-WILDS. CARE-Cell improves inflammation F1 over a UNI2-based cell classifier by 4.59, 8.35, and 1.84 percentage points on the three cell-classification datasets; under identical splits and GT instances, it improves over a CellViT++-based classification baseline by 11.95 points on Lizard and by a comparable margin on the private cohort. Substituting Virchow2 for UNI2 yields comparable gains, with UNI2 strongest overall. Without altering full-cohort predictions, the decision layer retains 92.15% of Lizard and 86.26% of PanNuke samples for automatic prediction, giving inflammation F1 of 0.9714 and 0.8765 on the accepted subsets, while hospital-shift stress tests illustrate failures outside the calibration assumptions. Together these separate three questions often conflated in expert augmentation: whether corrective evidence exists, whether its utility is identifiable, and which deployment decisions the available information can justify.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.