acceptodds
Under review as a conference paper at ICLR 2027

Compress Now, Contaminate Later: One Quantizer for Every Huber Level

Abstract

In binary hypothesis testing with compressed observations, the encoder may have to be fixed before anyone knows a bound on the contamination. We study this setting where the decoder learns a common Huber contamination bound only after encoding, and can then choose its test but not undo the encoder's merges. For known finite laws, output symbols suffice to keep a constant fraction of robust Hellinger information at every feasible level, where is comparable to clean squared Hellinger distance. We also characterize the cost of using fewer symbols, which can be steep. On a geometric family with evidence scales, the best uniform sample overhead satisfies , even though a bit selected separately for each level has constant overhead. An adjustable-growth version of our construction attains this order. The exponential penalty applies to one identical memoryless channel where a known schedule of different one-bit encoders has overhead on the same family. Learning the encoder from clean observations is harder still. No finite training budget depending only on clean difficulty preserves every raw-feasible level with a fixed output budget. A fitted test does succeed above an estimation-error threshold, and independent calibration certifies tests through simultaneous moment bounds. Experiments measure the alphabet–information trade-off, compare designs that target clean information, and quantify what certification costs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.