SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extraction
Abstract
Joint Entity and Relation Extraction (JERE) is highly sensitive to training data quality, making data augmentation a natural way to improve generalization. However, existing augmentation methods often weaken entity relevance and disrupt the structured semantics of annotated triples, limiting their effectiveness for JERE. We propose Structured Semantic Data Augmentation (SSDAU), a task-specific augmentation framework that preserves annotation-compatible triple semantics under the original entity–relation schema. SSDAU segments text by entity labels, matches candidate spans using hybrid semantic similarity that combines contextualized embeddings with traditional similarity scores, and reconstructs new instances through structure-preserving replacement. To further ensure semantic faithfulness, it applies consistency-aware filtering with pair scoring and BERTopic-based topic control to remove incompatible augmentations. We evaluate SSDAU on datasets with different annotation types, comparing five representative JERE models against seven popular augmentation baselines. Results show that SSDAU consistently improves JERE performance across benchmark settings, produces structurally and semantically compatible augmented data, and is more robust to ambiguity than non-LLM methods, yielding a substantially smaller average relative F1 drop (8.93% vs. 21.94%). Our code and data are available at https://anonymous.4open.science/r/SSDAU-B345.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.