Large Language Models as Preprocessing Optimisers in Clinical Spectroscopy
Abstract
Clinical spectroscopy has emerged as a powerful tool across biomedical research, with applications spanning from disease diagnosis and screening, to biomarker detection, therapeutic monitoring and intraoperative assessment. However, despite substantial progress, the analysis of spectroscopic data remains a major bottleneck, with current workflows still heavily reliant on trial and error, expert intuition and prior experience. Here, we introduce a multi-agent, large language model (LLM) framework for automated optimisation of preprocessing pipelines for spectral data. Our system combines a memory-derived information set, or dossier, with LLM-based agents that reason over previously evaluated preprocessing pipelines in order to propose, rank, and iteratively refine preprocessing strategies by following a multi-reward rule. Unlike exhaustive, brute-force search over a parameter space, the agent constructs candidates sequentially, enabling more targeted exploration. We benchmark our approach across 18 scenarios on four public clinical datasets. Our agent matches the accuracy of exhaustive grid search while evaluating fewer than 4% as many pipelines, and outperforms sequential stepwise optimisation in 17 of 18 scenarios. Ablation across LLM backbones and dossier configurations further shows that access to optimisation history is a key contributor to performance. These results demonstrate the potential of history-informed agents to enable more automated, adaptive and reproducible preprocessing pipelines in spectroscopy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.