AeroProc-Bench: Distinct Citation and Safety Failures Create a Deployment Dilemma in Aviation Maintenance LLMs
Abstract
Large language models deployed as maintenance assistants can cite manual references correctly while silently omitting mandated safety steps, a failure invisible to citation-based evaluation. We introduce AeroProc-Bench, a 4-dimensional benchmark built from 14,721 real aircraft fault-isolation items, and use it to reveal a deployment dilemma: citation accuracy and safety coverage are structurally independent per item () yet covary across methods (), so no training-free method optimizes both. The dilemma manifests as a ranking inversion at the frontier: the safest model achieves SEF 0.834 but scores only CEF 0.196, meaning a citation-only deployer would reject the safest available model. Supervised fine-tuning escapes the dilemma by dominating both axes, but is unavailable for frontier APIs. Single-metric evaluation therefore systematically misleads deployment in safety-critical domains; multi-dimensional, per-domain validation is essential. Code and scoring scripts are released in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.