MARA-NIR: Masked Autoencoder Representation Analysis of Near-Infrared Spectra
Abstract
Near-infrared (NIR) spectroscopy predicts the chemical composition of a sample from a single one-dimensional spectrum. Before calibration, chemometricians preprocess this spectrum with fixed formulas, such as scatter corrections and derivative filters. These formulas tell us in advance what a useful representation of NIR spectra computes, which is rare in representation learning. We use them to ask what a masked autoencoder learns without labels. We pretrain one on 7.2 million NIR spectra of food and feed and compare its encoder with the formulas, with untrained encoders and with supervised models. Its first layer rediscovers Savitzky-Golay derivative filters, which neither untrained encoders nor supervised models trained on soil carbon learn. Deeper in the encoder, soil organic carbon becomes more predictable, while untrained encoders stay flat. We then test the frozen encoder on the public LUCAS soil library, a new domain measured on a new instrument. With a linear model, it matches tuned chemometric models for soil organic carbon when all labels are used, and with 50 to 400 labels it stays within 0.05 in of the best of them.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.