How Advice-Routing Updates Shape Error Recognition in Language Models
Abstract
Language models that can consult an advisor must decide which of their answers need help. We study how training a model to request advice changes its ability to rank incorrect answers above correct ones, which we call error recognition. Starting from models trained for confidence calibration, we isolate part of the learned routing update and intervene on the activity it produces. In Qwen, a distributed subset of the update improves error ranking when all model variants judge the same proposed answers. Replacing this activity throughout the prompt with fitted constants preserves roughly half of the partial router's improvement over its starting model. Restoring activity in a later group of selected layers recovers much of the lost performance, and activity from the same question gives greater recovery than activity from another question. This pattern repeats across two factual question-answering datasets and a second routing seed. The results reveal a substantial gain under constant internal additions and a smaller, reproducible benefit from question-specific restoration. They do not establish a dedicated correctness representation or a consistent advantage in advisor decisions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.