acceptodds
Under review as a conference paper at ICLR 2027

First-Order In-Context Shapley: Scalable Data Auditing for LLM Alignment and Evaluation

Abstract

Post-training alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality: massive preference and instruction-tuning corpora inevitably accumulate label noise, safety inversions, and inter-annotator contradictions that surface-level deduplication cannot detect and that are too costly to find by LLM-judging every record. While Data Shapley provides a principled framework for valuing individual records, exact computation requires intractable retraining, and fast embedding-space surrogates such as KNN-Shapley fail to separate high- from low-utility demonstrations ( with generative utility). We introduce an inference-only valuation pipeline that computes the first-order term of an in-context cooperative game over semantic -NN neighborhoods in forward passes without parameter updates, achieving positive rank correlation with multi-shot Monte-Carlo in-context Shapley on of targets (mean , close to the estimator's own split-half agreement of ) with less compute. By projecting pairwise log-likelihood shifts onto a directed influence graph, we characterize each record through three signals: outgoing Exemplifying Advantage, incoming Exemplified Advantage, and a topological One-of-a-Kind Score that protects rare tail records. Pre-filtering with these signals restricts arbitration to of records before routing each candidate and a contrasting twin to an LLM arbitrator, achieving enrichment in arbitrator-confirmed annotation errors on graded response-generation data (HelpSteer2) and on pairwise preference data (Anthropic HH-RLHF), including contradictions that zero-shot likelihood and margin filters miss. During DPO training, confirmed preference contradictions oppose the realized update far more often than clean records from early on ( vs. in exact LoRA gradients) and stay below chance in reward accuracy. Finally, auditing the HH-RLHF evaluation benchmark exposes label errors on which aligned models disagree with the benchmark of the time ( when three frontier arbitrator families unanimously agree), and retraining DPO from scratch after repairing or filtering the flagged training records improves reward accuracy by overall and on the repaired test inversions, even without the LLM arbitration step.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.