Reliability-Aware Determinantal Point Processes for Robust Informative Data Selection in Large Language Models
Abstract
Traditional data sampling methods, including those based on Determinantal Point Processes (DPPs), offer deterministic approaches to maximize diversity, assuming that the selected data batches are always available without error. This presumption prohibits their use under probabilistic and disrupted access due to storage outage, imperfect communication, and hardware failures. This problem arises naturally in real-world LLM systems that aggregate data from distributed sources (e.g., RAG pipelines, teacher-student fine-tuning, and LLM-based cooperative perception and planning). We show that the original formulation of DPP-based inference collapses under such conditions. To address this gap, we introduce ProbDPP, a novel reliability-aware implementation of k-DPP that accounts for probabilistic data access by regularizing and decomposing the objective function into a geometric diversity term and unreliability cost, which remains well-posed, facilitating robust batch selection under uncertainty. To cover cases where failure probabilities are not known in advance, we frame this reliability-aware diversity maximization as a combinatorial semi-bandit problem and propose an algorithm that naturally adopts an efficient online learning policy. We evaluate the proposed method on different tasks, including question answering, RAG pipelines, summarization, and image reconstruction, across different benchmarks. Theoretical analysis provides regret bounds for the proposed approach under uncorrelated access links, ensuring performance guarantees.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.