Multi-task Prediction of Solid-phase Extraction Protocol Components from Molecular Descriptors
Abstract
Solid-phase extraction (SPE) is widely used to enrich trace pollutants before instrumental analysis, yet selecting cartridge chemistries, solvents and mixture ratios still relies on literature search and expert adaptation. Learning these choices from published methods is difficult because a pollutant can have several reported alternatives, solvent labels are highly imbalanced, and ratio annotations are incomplete. We propose SPEGen (Solid-Phase Extraction Protocol Generator), a multi-task framework that predicts SPE components from molecular descriptors. We aggregate published methods by pollutant to construct set-valued supervision that retains alternative reported choices. A shared encoder with task-specific heads predicts cartridge sets, workflow-step indicators, step-specific solvent sets and their ratios, with ratio training restricted to observed annotations. Rank-Gauss normalisation handles heavy-tailed descriptors, while a validation-tuned decoder selects label sets and suppresses solvents for absent steps. Bootstrap aggregation combines predictions from resampled training data. Experiments against six tabular baselines on a held-out pollutant split show that SPEGen-Bag achieves the highest cartridge and solvent F1 and the lowest ratio mean absolute error. Cumulative ablations show gains from normalisation and bagging, while decoding analyses reveal how predicted set size affects the trade-off between recovering reported alternatives and exact-match accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.