Fidelity Probes for Specification-Code Alignment
Abstract
Software modernization often relies on natural-language specifications recovered from an existing implementation. Incorrect or missing requirements can lead to defects in the modernized system. We introduce *fidelity probes* to check a specification against the code and guide its revision. Each probe pairs a question about program behaviour with a reference answer derived from the code. A language model answers the question using only the specification, and the fraction of agreeing answers defines *fidelity*. Disagreements flag conflicting claims or missing information and guide proposed corrections and additions to the requirements. An LLM generates probes directly from code or phrases facts selected through control-flow, data-flow, and call-graph analysis. We use an observability rule to focus probe questions on behaviour that the modernized system is intended to preserve, including user-visible outputs, changes to stored business data, and interactions with other systems. We revise the specification using fresh probes and evaluate each version on a fixed held-out set. A repair–regression model describes how new failures can offset the gains from repairs. On CardDemo, held-out fidelity improves from 0.59 to 0.89, exceeding free-form revision from source code. We also evaluate the method on industrial systems and public documentation, and assess probe quality and reported defects through human reviews and comparison with an independent requirements audit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.