acceptodds
Under review as a conference paper at ICLR 2027

GNOSIS: Consistency-aware Probes for Selective Prediction over Internal States

Abstract

Selective prediction determines whether an LLM's output should be trusted or deferred, which is critical for building reliable LLM systems. Existing confidence estimators fall into three groups: multi-sample methods requiring repeated generation, post-generation scoring over fully decoded outputs, and pre-generation hidden-state probes. While pre-generation probes yield confidence estimation before token generation and avoid decoding overhead, existing approaches rely on a narrow readout paradigm: they typically preselect a single readout layer and optimize against only one training objective, limiting their overall performance. We present GNOSIS, a suite of lightweight single-pass pre-generation probes which operates on the prompt-final hidden state of a frozen LLM. It comprises three core designs: (1) The Direct Correctness Probe (DCP) infers answer correctness and adaptively selects the optimal readout layer via cross-validation instead of manual specification. (2) Anchor-Regression Distillation (ARD) distills multi-dimensional cross‑model consistency signals from a stronger teacher model into student hidden states, where the teacher is only used for target construction and is never invoked during inference. (3) A Fusion module jointly optimizes both probe heads on a shared trunk to exploit complementary error patterns. Across 11 LongBench tasks spanning eight base models (3B-72B), GNOSIS achieves competitive or state‑of‑the‑art AUROC and PRR. When deployed as an out‑of‑distribution HotpotQA cascade router, GNOSIS allows escalated queries to fully bypass student‑model generation, substantially lowering latency, token consumption, and API cost. Ablations confirm that DCP and ARD capture distinct failure modes and the joint multi‑objective Fusion outperforms naive post‑hoc score averaging. Our results demonstrate that adaptive readout-layer selection together with multi-faceted training targets is essential for high-performance pre-generation selective prediction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.