Same Text, Different Prediction: Serving-Context Nondeterminism in Text Classifiers
Abstract
Deterministic inference is essential for reliable and trustworthy machine learning. Prior studies of text generation have shown that changing factors such as batch size, batch composition, hardware, or inference engine can substantially alter the generated text, even when the prompt, model parameters, and sampling randomness are fixed. These differences have been attributed in part to floating-point non-associativity, shape-dependent kernel selection, and other implementation-level differences in numerical execution. However, it remains unclear whether, when, and to what extent the same factors affect text classification. To the best of our knowledge, this is the first systematic study of serving-context non-invariance in text classifiers. We train 180 models (spanning discriminative, pseudo-generative, and fully generative classifier formulations), and evaluate them across four different categories of serving contexts. We show that label stability can conceal substantial score instability e.g. changing only the batch shape changes no labels across fp32 comparisons, yet under bf16 it redistributes up to 56.7 percentage points of predicted probability mass, with label flipping concentrated around small margins. Fully generative classifiers are particularly susceptible, compared to their discriminative counterpart. We derive sufficient conditions for label stability under various serving contexts and give a separate mitigation for each non-invariance mechanism as well. Taken together, our results identify and quantify the previously understudied but important serving conditions that must be factored in for reproducible text classification.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.