acceptodds
Under review as a conference paper at ICLR 2027

RFWall: A Structural Diagnostic of Generative Data via the Augmentation Difficulty Index

Abstract

Generative data augmentation is widely used to reduce labeling cost, and RF fingerprint localization, where every label requires a site survey, has become a prominent application. Reported gains are hard to compare: evaluations rarely share data or splits, focus on dense Wi-Fi, and often fit the generator on data withheld from the low-label condition. We introduce RFWall, a benchmark of 31 augmentation methods on 20 partitions spanning Wi-Fi, cellular, and non-RF controls, under a single leakage-controlled protocol. Fitting generators on the full labeled pools inflates the measured gain at a label budget by . Under leakage control, augmentation recovers of the available accuracy headroom on Wi-Fi and none on cellular fingerprints, and its positive effects under a nearest-neighbor localizer reverse under an MLP. Grounded by the Data Processing Inequality (DPI) on information bounds, we conduct controlled sparsification experiments on high-density physical testbeds and identify representational feature vocabulary as the causal driver of this gap: shrinking the vocabulary from 520 columns to 17 elevates exact label collision from 0.004 to 0.853, collapsing manifold structure and reducing augmentation gains to zero. We distill these mechanisms into the Augmentation Difficulty Index (ADI), a data-inherent screen combining label collision, relational manifold sparsity, and noise level, which orders held-out dataset utility at Spearman under leave-one-dataset-out validation. Beyond the RF domain, spatial voxel experiments on non-RF controls show identical label collision dynamics. Controlled noise injection confirms ADI's theoretical bounds, with structure-aware generators outperforming baseline models across all test cells. Our findings demonstrate that simple non-parametric interpolation frequently matches or exceeds deep generative models, and that standard distributional fidelity metrics fail to predict downstream utility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.