acceptodds
Under review as a conference paper at ICLR 2027

EMPIRE: Explicit Method for dataset Partitioning using Integer programming for Named Entity Recognition to reduce data leakage

Abstract

Evaluations of Named Entity Recognition (NER) models are often confounded by data leakage between dataset splits. Entity leakage occurs when the same entity appears in both training and test data, and context leakage when test sentences share their surrounding context with training sentences; both let models score well through memorization rather than generalization. We introduce EMPIRE (Explicit Method for dataset Partitioning using Integer programming for Entity Recognition), which formulates dataset splitting as an integer linear program. It reduces entity and context leakage, weighted by a single parameter , while fixing the split sizes and bounding each entity class's share of every split. On 24 NER datasets, EMPIRE matches the entity disjointness of a minimum-cut baseline while preserving class balance, and with it reduces context leakage more than a similarity-based baseline. Removing entity leakage lowers the test F1 of both a BiLSTM-CRF and a fine-tuned Llama-2-7B, whereas removing context leakage alone does not, indicating that much of the F1 on native splits reflects entities seen during training. The class-balance constraint is essential for this comparison: without it, splits strand entity classes and lower F1 for reasons unrelated to generalization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.