The Vanishing Middle: Reinforcement Learning Erases the Uncertain Category in Ordinal Classification
Abstract
In high-stakes decisions, such as clinical diagnosis or loan approval, the most useful answer is often “we cannot decide”: an uncertain middle category that flags the case for a human and breaks the ambiguity. We show that reinforcement learning with verifiable rewards (RLVR), a standard way to improve large language model classifiers, systematically erases exactly this category. The setting is ordinal classification, whose labels lie on an ordered scale with a distinct middle. In detail, RLVR pushes answers toward the most common categories, and the rare middle loses recall. However, overall accuracy actually rises, making it hard to notice that the “we cannot decide” output is gone. Rebalancing the classes is not a cure we can deploy, and a second cause is at work: the middle is a two-sided range (two cutoffs), and at a matched frequency it collapses far more than a one-sided end (for example, a confidently positive or negative label). The fix therefore has to be structural. We propose CURB (Cumulative-ray Reparameterization of the Bounded middle), which rewrites the middle as two one-sided cutoff decisions at the model’s output. CURB restores the middle across every setting and all three model families, without making it dominant.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.