acceptodds
Under review as a conference paper at ICLR 2027

Certified Selective Gates for Video-Language Models Exceed Their Error Target Before Genre Drift Begins

Abstract

A selective gate lets a frozen video-language model answer only when its error rate on answered questions is certified below a target α, and the standard repair when the genre mix drifts is to detect the break and recalibrate on fresh labels. We find that on most drift streams there is nothing for the repair to restore. A split-conformal gate certifies the false-accept rate, and selective error equals that rate times the base error rate divided by coverage, so at the certified level the gate is over α wherever coverage falls below the base error rate, before any drift begins. We test this on 5 backbones (3B–32B) and 4 video-QA benchmarks with a per-genre certified gate, a label-free coverage-drift monitor, and drift streams built from benchmark metadata. The frozen gate is already over α before the break on 30 of 35 stream cells. Where α did hold before the break, recalibration on a deployment-sized label budget restores it on none, an oracle that recalibrates at the true break with no delay does no better, and a pre-registered sweep to a larger budget shows that more labels help but do not restore the target; the holds that do occur are thin under question-clustered upper bounds. A rule that certifies selective error directly is the honest alternative, and under a Bonferroni grid it is vacuous in most realizations at either budget, answering a third as often as the conformal rule. The quantity to certify, and the coverage regime in which certifying it is possible, are fixed by the identity before a label is collected.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.