acceptodds
Under review as a conference paper at ICLR 2027

Characterizing Errors in Small Medical Language Models: Geometry, Prediction, and Causal Sensitivity

Abstract

What do differences between correct and incorrect medical answers reveal about a language model's errors? Geometric separation, predictive internal signals, and intervention effects offer different evidence, yet none alone establishes that a model recognizes or can correct its mistakes. We introduce ***Mirage***, a framework that tests these claims separately on shared questions and responses, using coverage checks, predictive baselines, and controlled interventions. Across five medical checkpoints and four datasets, we collect 208,896 responses across calibration and evaluation. Incorrect responses are more dispersed in 545 of 808 eligible evaluation pairs, but this geometric comparison covers only 12.6% of the 6,400 evaluated pairs. On 278 matched pairs, masking explicit answers reduces biomedical Fisher label propagation macro-F1 from 0.807 to 0.670, revealing a substantial contribution from answer content. Internal states predict errors on unseen questions, while gains beyond answer type vary across visual tasks and fitting partitions. Interventions produce protocol dependent effects, and a supplementary restoration study finds no observed benefit in sampled answer recovery. Together, these results identify three obstacles to interpreting correctness signals: selective coverage, dependence on answer content and task structure, and a gap between output sensitivity and demonstrated repair. By turning these obstacles into explicit testbeds, we provide a reproducible framework for testing what geometric, predictive, and intervention evidence can establish about model errors, while keeping benchmark agreement separate from clinical explanation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.