acceptodds
Under review as a conference paper at ICLR 2027

Separating the Model from the Decision Rule in the Evaluation of Long-Tailed ICD Coding

Abstract

Automated ICD coding is extreme multi-label classification over thousands of long-tailed codes; most reported progress since 2023 is in macro-F1 and has been credited to new model components. Macro-F1, however, also depends on the decision rule, the thresholds that turn a model's scores into codes. Published systems on the standard MIMIC-IV ICD-10 benchmarks differ in how they set these thresholds or do not state how, so the model's share of any reported gain is unknown. Changed on their own, the thresholds can add as much macro-F1 as nearly half of that progress and alter how systems compare; we therefore propose a reporting standard that separates the model from the rule. The benchmark's reference protocol tunes one threshold on validation for micro-F1; using one per training-frequency bin, tuned for twice micro-F1 plus macro-F1, instead raises the macro-F1 of PLM-ICD, the 2023 state of the art, by +0.034, 46% of the progress since. On the reproduced members of the state of the art, GoM-ICD, the same change lifts one member alone above the published three-member ensemble on macro-F1. The gain is positive on all of our more than 60 training runs and all five released checkpoints we scored. The reason is that rare codes keep a high per-label AUC while a threshold shared with common codes sits above their positives: they are rankable but not decidable. Pre-registered tests on legal and protein-function benchmarks find the same, and the gain grows with how well the tail is ranked. Reporting both rules also shows which components improve the model: averaging members that rank codes differently gains under both, while negative focusing in the loss only moves the operating point. Built from the components that gain under both, an average of three long-context clinical encoders sets the state of the art on both standard benchmarks. On the 7,942-code benchmark it improves on GoM-ICD by +0.057 macro-F1 at higher micro-F1, of which +0.048 comes from the thresholds; on the 26,096-code benchmark it reaches 1.4x the best published macro-F1 at equal micro-F1. The proposed standard reports every result under both rules; a difference that holds under both is a difference between models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.