acceptodds
Under review as a conference paper at ICLR 2027

Conformal Maps for LLM Evaluation: Learning Behavior-Aware Instruction Geometry

Abstract

Aggregate benchmark scores collapse heterogeneous patterns of model success and failure into a single number. A more informative evaluation should characterize this variation at the level of individual instructions, relating a held-out instruction to previously observed model behavior while quantifying uncertainty in that prediction. Conformal prediction provides a natural framework for the latter through finite-sample coverage guarantees, but the informativeness of its prediction sets depends on the score being calibrated. We introduce *conformal maps*, learned instruction embedding spaces whose local support is optimized for informative label-conditional conformal prediction while remaining anchored to a pretrained semantic representation. The same measure of local support organizes historical successes and failures, defines the conformity score, and exposes the evidence surrounding a new instruction. Across five QA and reasoning benchmarks and three target models, the learned maps induce measurable behavioral organization, retain a controllable amount of pretrained structure, and improve conformal singleton rate over the pretrained representation in most dataset–model pairs while maintaining label-conditional coverage. The learned geometry also exposes the amount, direction, and provenance of local behavioral evidence, providing an inspectable account of which prior instructions support a prediction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.