acceptodds
Under review as a conference paper at ICLR 2027

Dynamic Reliability Curriculum Guided Reinforcement Learning for Partial Label Disambiguation

Abstract

\em Partial-label learning (PLL) addresses the problem where each training sample is associated with a set of candidate labels, among which only one is the ground-truth label. The ambiguity of candidate labels brings great uncertainty to model optimization. Traditional disambiguation methods usually rely on classifier confidence or predefined label refinement rules. Such strategies may be affected by confirmation bias, leading to error accumulation and label noise propagation. To address these issues, this paper proposes a \em dynamic reliability curriculum guided reinforcement learning framework for PLL, called DC-RPLL. This method models label disambiguation as a Markov decision process, independently deciding for each sample whether to \em update, \em retain, or \em abstain. Specifically, DR-CPLD first represents each candidate label set as a soft label distribution, jointly uses classifier predictions and local feature space structure to construct disambiguation labels. Meanwhile, the reliability of each sample is estimated from neighborhood consistency, feature affinity, and soft label consistency to quantify the credibility of local structures. Based on this reliability, this paper designs dynamic curriculum thresholds so that highly reliable samples enter policy updates early, while low-reliability samples are gradually added later, thereby alleviating early error propagation. The Actor-Critic policy then guides policy learning according to the current confidence gain, neighborhood reliability, and label drift penalty, i.e., \em update, \em retain, and \em abstain. In addition, we develop a reliability-aware curriculum learning mechanism that gradually accepts samples from high to low reliability during training, effectively suppressing early error propagation. Experiments on benchmark and real datasets verify the effectiveness of the proposed algorithm.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.