acceptodds
Under review as a conference paper at ICLR 2027

Access and Representation Barriers to Auditing under Proxy-Score Optimization

Abstract

Language models are optimized against cheap proxy scores, but how costly is it to audit the resulting target change? We identify access and representation barriers linked by likelihood-ratio geometry. Baseline archives face a no-hit lower bound, while active auditors trade observable-transcript KL against sample count. An exact duality characterizes worst-case target change hidden from a fixed feature span. For exact single-prompt one-sided exponential-power tilts, minimax log archive complexity and half-mass likelihood-ratio-tail rarity asymptotically match optimization KL. Using exact conditional reweighting on a finite pool of judge-scored Qwen2.5-7B math traces, we find a -order increase in the tested archive requirement, seven practical samplers above the movement–sample benchmark, and substantial held-out error in fitted score-blind monitors. At 18 score-matchable Group Relative Policy Optimization (GRPO) checkpoints, the plug-in KL residual from a score-matched one-parameter tilt accounts for – of measured KL, indicating substantial departure from that tilt path. Reported LR-tail estimates differ substantially from score-tail rarity, although differing likelihood-evaluation paths and a finite reference pool limit their interpretation. This motivates measuring LR-tail overlap under a common likelihood convention before using proxy-score tails as overlap diagnostics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.