acceptodds
Under review as a conference paper at ICLR 2027

MedShard: Task-Structured Inference Improves Clinical Quality Control across Model Scales

Abstract

Structured clinical-document audits require a language model to track many rules, locate their evidence, and emit a variable number of findings. We introduce MedShard, a task-structured inference harness that compiles this workload into rule-specific execution paths. MedShard preserves rule-relevant fields, resolves eligible cases with deterministic prechecks, and combines focused model judgments with a deployed regex system through development-calibrated per-rule routing. On 250 Chinese admission notes evaluated against 19 audit rules, MedShard with a fixed 27B model achieves micro-F1 of 0.751 (95% CI [0.726, 0.776]), compared with 0.355 for direct prompting and 0.621 for the regex system. The complete LLM tier reaches 0.702, and calibrated routing improves on blanket union by 0.016 F1. Across 4B, 9B, 27B, and 122B-A10B configurations, MedShard improves F1 by 0.379–0.532 over each checkpoint's direct-prompt baseline. The routed 4B system exceeds direct 122B-A10B, while dense 27B achieves the highest routed score. On two mutually exclusive error-type benchmarks, decomposition increases detection precision while reducing type accuracy. These results connect rule-level inference design, evidence-source complementarity, and output structure to the performance of locally deployed clinical-document quality control.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.