What to Learn for Sequential Log-likelihood Ratio Tests?
Abstract
In a learned log-likelihood ratio (LLR) test, one checkpoint, or a model state saved during training, must be selected to generate the LLR estimates used for threshold-based decisions. However, we find that conventional validation criteria can favor checkpoints whose estimated LLR trajectories poorly match the shape of the true LLR. Moreover, multiplying both the LLR and thresholds by the same positive constant leaves stopping times and decisions unchanged, thus test results alone cannot determine exact-scale LLR estimation error. To establish a proper validation metric for checkpoint selection, we therefore shift the estimand to the positive-projectivized LLR (P+LLR), the true LLR up to positive rescaling. With correspondingly rescaled decision thresholds, P+LLR gives exactly the same stopping decisions as the true LLR to preserve optimality guarantees of statistical tests. P+LLR lets us introduce a validation metric Normalized Integrated Tradeoff Area (NITA), which is provably ordered with an oracle P+LLR estimation error in a high-threshold detection task under an i.i.d. data stream with independent noise. We also establish a finite-sample checkpoint-selection guarantee for empirical NITA and provide a unified extension of the theory to broader non-i.i.d., multiclass data and classification tasks, accompanied by experiments on 12 diverse datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.