acceptodds
Under review as a conference paper at ICLR 2027

Capability Self-Assessment in Large Language Models

Abstract

The ability to recognize one’s own limitations and decide whether to solve a problem or seek help is fundamental to reliable intelligent systems. Yet modern large language models tend to overestimate their competence and attempt queries they cannot solve. We study how models can learn Capability Self-Assessment (CSA) while preserving their problem-solving ability. Because the same model both assesses and solves queries, self-assessment training can change the capability it must recognize. We systematically compare supervised fine-tuning (SFT) and reinforcement learning (RL), evaluating each trained model’s self-assessment against its updated capability and measuring how well it retains its solving ability. Across model families, scales, and domains, RL provides the strongest assessment–retention trade-off. We also find that CSA transfers to the out-of-distribution settings we evaluate, and demonstrate its practical value for local–cloud routing at inference time and targeted data selection during training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.