InvarFlow: Shortcut-Invariant Pre-training for Generalizable Encrypted Traffic Classification
Abstract
Pre-trained models on raw traffic bytes report encrypted traffic classification (NTC) accuracies above 98% on mainstream benchmarks, yet recent evidence attributes much of this progress to shortcut learning: models exploit dataset-specific artifacts that correlate with labels in collected data but carry no application semantics, and degrade under distribution shift, field obfuscation, and adaptive evasion. Existing mitigations either remove a pre-defined list of suspect fields—missing dataset-specific shortcuts—or diagnose shortcut reliance post hoc without feeding it back into training. We present InvarFlow, a pre-training framework that encodes shortcut invariance directly into the training objective, combining per-corpus shortcut profiling via adjusted mutual information, a differentiable suppression term built on a mutual-information upper bound, a cross-environment invariant-risk penalty, and a shortcut-aware masking curriculum. We formalize shortcut invariance as a constrained optimization problem and prove that, under an explicit semantic/shortcut decomposition, any representation achieving exact invariance attains the environment-robust optimum—and that approximate invariance bounds the cross-environment risk gap linearly in the residual dependence. A prerequisite gate experiment on synthetic ground truth shows that the deployed mutual-information upper bound is loose but directionally valid (monotone, above-truth), which we turn into an explicit design rule: the bound drives training while calibrated probe statistics drive evaluation. In a controlled toy study with an injected, abolished-at-test shortcut, the invariance objective reduces the cross-environment accuracy gap by 71% relative to ERM at a 3-point in-distribution cost, matching the trend predicted by our bounds. On a large-scale benchmark under our pre-registered protocol, InvarFlow cuts the cross-environment gap by 72–74% relative to ERM while staying within 0.4–0.9 F1 of the best published in-distribution result per dataset, and its accuracy survives field-level decontamination with only 1.2–2.0 points of degradation—evidence that the gains come from genuine semantic reliance rather than artifact fitting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.