Diagnose Before Adapting: Intervention Headroom under Clinical Drift
Abstract
Retrieval systems can adapt their evidence memory, decision rule, or model parameters, yet these interventions are often evaluated separately. We study them jointly through Diagnose-Before-Adapt (DBA), a prequential intervention audit that measures conditional gains under delayed label release. Recent versus frozen memory measures staleness, cumulative memory measures retention, a support-constrained oracle bounds candidate-restricted decisions, and learned aggregation measures realized gain. Together, these probes characterize the interactions among adaptation components. We apply DBA to 117,328 consultations over 51 months and introduce Learned Neighbourhood Aggregation (LNA), a lightweight ranker over prior evidence with frozen text encoders. Fresh-memory headroom grows by more than five balanced-accuracy points per year, while retaining labelled history recovers 20.33–34.98 points. Against recalibrated cumulative memory, base LNA gains 2.83, 1.05, and 3.53 points on the three streams, with over lower reported GPU occupancy than monthly retraining. Gains are positive in 44 of 45 correlated seed–origin conditions and under label-blind overlap filtering. Public TCM-SD further replicates the aggregation design on a static task. Together, these results identify memory retention and learned evidence use as effective, resource-efficient adaptation components in the studied longitudinal streams.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.