acceptodds
Under review as a conference paper at ICLR 2027

Matched Parent-Removal Audits Expose When Augmentation Provenance Changes Decisions

Abstract

Failures in augmented training corpora propagate from original assets through rewriting, filtering, distillation, and merge operations, making the root assets that require inspection or removal difficult to identify. We introduce a matched parent-removal audit that isolates the decision value of provenance by fixing the target model, candidate set, intervention, and top- action while varying the information available to the ranker. Across four logged text and code workflows, graph structure improves top- recovery over flat attribution, while a learned path ranker adds – points over structure-only Metadata Provenance. The ordering persists under hard-negative and reachable-only controls built from the same candidate manifests. Controlled operator-label corruption further maps where learned structured rankers degrade relative to structure-only provenance. Together, these results establish fixed-decision auditing as a rigorous way to measure the decision value of recorded provenance for parent ranking.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.