Criteria as Code: When Written Rules Beat Learned Ones, and When They Don't
Abstract
Written clinical criteria offer an explicit alternative to learning every diagnostic decision from annotated images. But when does protocol-derived decision logic transfer better than logic learned from data? We study this question with CriteriaCode, which separates multiple sclerosis (MS) lesion segmentation into learned candidate detection, 34 radiological attributes, and an executable criteria program with calibrated thresholds. The main comparison uses a hand-written program and a shallow decision tree trained on the identical clinically informed attributes. Both expose their decision logic, while its source differs: the annotation protocol or the training labels. In-domain, the tree exceeds the program from 32 annotated patients, and the end-to-end network remains more accurate. Under transfer to MSLesSeg, the program achieves lesion Dice of 0.587 versus the tree's 0.550 with FLAIR input, matching the source-trained network's mean Dice. On Shifts, the ordering reverses. Controlled noise and bias-field experiments do not reproduce this reversal as a function of severity alone. Attribute analysis instead identifies greater orientation and multi-contrast drift on Shifts. Removing the associated rules and recalibrating closes 70% of the FLAIR program-tree gap there, but reduces multi-contrast performance on MSLesSeg. This intervention supports an attribute-specific explanation rather than a general robustness advantage for written rules. A threshold sweep also exposes a mismatch between a literal 3 mm size rule and its implementation on thick-slice images. LLM-generated programs are less reliable on real candidates than the hand-written reference. Together, these results clarify the benefits and limits of protocol-derived structure: its transfer value depends on the attributes affected by the target shift, while explicit rules make those dependencies inspectable and editable. We also correct two label-orientation errors in the public MS3SEG release; anonymized code and every generated program accompany this submission.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.