TAPS: Tree-Adaptive Prediction Sets via LCA-Weighted Nonconformity Scoring for Hierarchical Multi-Label Clinical Coding
Abstract
Conformal prediction (CP) provides distribution-free, finite-sample coverage guarantees. However, applying CP to hierarchical multi-label classification leads to a combinatorial explosion over possible label subsets. Recent parallel work (Mortier et al., 2026; den Hengst et al., 2025) addresses this by allowing prediction sets to include internal nodes of a taxonomy. We identify a complementary gap: prior methods either use flat leaf-level conformity scores or incorporate hierarchy without an explicit abstraction-cost penalty in the conformity score. We introduce **Tree-Adaptive Prediction Sets (TAPS)**, a CP algorithm that directly integrates hierarchical structure into the calibration score. TAPS employs a *specificity-regularized Lowest-Common-Ancestor (LCA) score*, which penalizes each tree node based on its uncertainty and a height-based abstraction cost (weighted by ). Crucially, this identical score governs both calibration and inference. We prove that the TAPS score is dominated by the flat score, ensuring calibration thresholds remain no larger than those of flat CP at the same . We also prove TAPS guarantees marginal semantic coverage under exchangeability, with calibration and inference provably evaluating the exact same event. The parameter controls an empirical compactness-specificity trade-off while conformal validity is preserved after independent selection, requiring a non-degeneracy screen to prevent trivial root-collapse. On MIMIC-IV automated ICD-10 coding (QLoRA-tuned Phi-3-mini, full training set), TAPS attains near-nominal empirical semantic joint coverage, comparable to Flat LAC, while producing prediction sets roughly 48% smaller than flat CP. It achieves this by summarising uncertainty via ancestor nodes, trading off exact leaf-level specificity. We evaluate against eleven baselines spanning flat, hierarchical, and multi-label conformal methods, and identify a common failure mode, which we call the 'free root' pathology, that causes every hierarchy-aware baseline we test lacking score-level integration to exhibit severe undercoverage, root-collapse, or substantially inflated prediction sets. We further validate TAPS under the same calibration and evaluation framework on a second, structurally distinct domain, the arXiv category taxonomy (depth , vs. ICD-10's depth ), providing evidence that both the compactness gains and the baseline failure modes are not specific to clinical coding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.