Evidence, Adaptation, and Tail Risk in Verifiable-Reward Taxonomy Alignment
Abstract
We audit verifiable-reward and supervised post-training for institutional taxonomy alignment: mapping national occupational records into ISCO-08 across 878 country/entity–system–source cells. On a development-exposed benchmark, five row-only post-SFT variants leave worst-cell reliability within three points of SFT, and five checkpoints share a 73-cell worst-decile core. Continued SFT on evidence-augmented prompts with gold targets only (three seeds) changes this pattern: worst-cell conditional value-at-risk increases by 0.072–0.081 relative to plain-prompt SFT, while overall strict accuracy falls by 4.7 points to 74.2–74.4%. The change is heterogeneous: cells defined as difficult from development records improve from 14.3% to 24.1–24.2%, while easy cells fall from 99.3% to 91.3–91.6%. At the coverage of the explicit ABSTAIN-target control, confidence-selected gold-only SFT reaches 96.1–96.2% selective accuracy versus 84.0–84.3%; at the strongest selective RL policy's 70.8% coverage, it reaches 92.6–92.8% versus 63.3%. On a prospectively frozen, single-source transfer of 976 queries, five labeled references raise gold-only accuracy from 9.1–9.9% to 28.9–29.6% at fixed weights; below a post-hoc nearest-code lookup (35.8%). These results identify recipe-specific tradeoffs between institutional-tail accuracy, aggregate error and deferral. They do not establish a general information-theoretic limit of verifiable rewards, or isolate training evidence from the additional optimization in continued SFT.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.