Leveraging Hard Negative Samples: Weakness-driven Policy Optimization for Generative Information Extraction
Abstract
While Large Language Models (LLMs) show remarkable capabilities in Generative Information Extraction, their performance remains constrained by the uncontrollable nature of pre-defined label generation. Existing research, centred on prompt engineering or supervised fine-tuning, offers limited gains due to the neglect of hard negative samples. However, leveraging these samples introduces challenges, including deficient error-type awareness, inadequate training focus, and weak dynamic adaptability. To address these issues, we propose a novel framework titled Weakness-driven Policy Optimization (WPO). It first collects error samples that reveal LLM weaknesses, followed by pre-defined atomic error types. Then, WPO uses an LLM to mine composite error types to gain an error distribution. Subsequently, the framework constructs a training set via error-centric data synthesis and error-type-based data augmentation. Finally, WPO employs a dynamic weight allocation mechanism that strategically assigns weights to original training, synthetic, and augmented samples, fine-tuning the LLM using Group Relative Policy Optimization (GRPO). Experimental results on the Qwen3 series LLMs show that WPO achieves a performance gain of up to 1.12% in F1 and 5.31% in precision on Qwen3-8B in Relation Extraction compared to the state-of-the-art (SoTA) baseline. Furthermore, extensive analysis validates the cross-task generalization and cross-scenario robustness of our method.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.