acceptodds
Under review as a conference paper at ICLR 2027

AeroProc-Bench: Distinct Citation and Safety Failures Create a Deployment Dilemma in Aviation Maintenance LLMs

Abstract

Large language models deployed as maintenance assistants can cite manual references correctly while silently omitting mandated safety steps, a failure invisible to citation-based evaluation. We introduce AeroProc-Bench, a 4-dimensional benchmark built from 14,721 real aircraft fault-isolation items, and use it to reveal a deployment dilemma: citation accuracy and safety coverage are structurally independent per item () yet covary across methods (), so no training-free method optimizes both. The dilemma manifests as a ranking inversion at the frontier: the safest model achieves SEF 0.834 but scores only CEF 0.196, meaning a citation-only deployer would reject the safest available model. Supervised fine-tuning escapes the dilemma by dominating both axes, but is unavailable for frontier APIs. Single-metric evaluation therefore systematically misleads deployment in safety-critical domains; multi-dimensional, per-domain validation is essential. Code and scoring scripts are released in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.