Code, Not Chains-of-Thought: Localizing Guideline-Adherence Risk by Compiling Guidelines into MedCode
Abstract
Clinical large language models can cite guidelines fluently while producing unsupported dosages, staging criteria, or treatment recommendations. We argue that this failure is partly architectural: when the same model both interprets evidence and emits the clinical action, there is no machine-checkable contract between a recommendation and the source guideline. We present MedCode, a model-agnostic framework that compiles oncology guidelines into executable, citation-bound Python rules and restricts the LLM to structured patient-state extraction. The rule engine, rather than the decoder, maps typed evidence to a closed ClinicalAction space with bounded verbatim citations and explicit abstention when required evidence is missing. This design does not eliminate clinical error; residual failures can still arise from extraction or rule authoring. Instead, it localizes guideline-adherence risk to two auditable choke-points, enforced by six static compliance gates and localized rule diffs under guideline updates. We instantiate MedCode on oncology staging and treatment tasks using NCCN-derived rule scaffolds and real clinical cohorts. On TNM staging, deterministic execution substantially improves smaller and open-weight models: Qwen3.6-27B rises from 25.95 to 54.91 TNM accuracy on LungCURE-EN, and Gemma-4-E4B-it rises from 17.31 to 58.70 on EsophagusPKUCH. On treatment recommendation, MedCode improves coverage of physician-specified gold-plan items across all evaluated models, although this rubric measures coverage rather than harmful or extraneous recommendations. Finally, the same executable rules serve as a verifiable GRPO reward for MedCode-LM: distilling a Qwen3.6-27B teacher into a Qwen3.5-4B student raises TNM accuracy from 35.64 to 50.34, approaching the teacher's 52.67 while preserving auditable decision semantics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.