An Evidence-Aligned Framework for In-Context Chinese Legal Named Entity Recognition
Abstract
Named entity recognition (NER) is a foundational information extraction task, and it must recover exact entity spans and types from limited demonstrations in the few-shot in-context settings. Chinese legal NER is particularly challenging because implicit word boundaries hinder exact span localization, while similar text spans may belong to different entity types depending on context. Existing methods often rely on whole-sequence or word-level retrieval, generate detached mention lists, creating a mismatch between the evidence provided and the evidence required for boundary- and role-sensitive decisions. In this paper, we propose C-NER, an in-context learning framework with a frozen LLM and no task-specific gradient updates. It retrieves demonstrations that contain local cues for locating entity boundaries, preserves predicted spans at their exact positions in the original text throughout extraction, and uses surrounding context to distinguish spans that may belong to different entity types. Experiments on four NER benchmarks, including complete-system and controlled-retrieval comparisons, show that C-NER achieves 90.99 F1 on CAIL2021 and 91.56 F1 on MSRA, outperforming the strongest reproduced baseline DEER by 3.65 and 1.87 F1 points, respectively, while ranking second on CLUENER2020 and OntoNotes. On CAIL2021, it uses one LLM call per instance and reduces LLM calls by 55.1% and token usage by 45.3% relative to DEER. Code, configurations, and reproducibility artifacts are available at https://anonymous.4open.science/r/charcon-ner.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.