acceptodds
Under review as a conference paper at ICLR 2027

SDA: Separating Discovery from Annotation for LLM-Based Structured Extraction

Abstract

Structured extraction organizes semantic elements and their relations in free-form text into predefined records. Large language models (LLMs) support this task through in-context learning (ICL), yet often fall short of task-specific supervised models. Prior work addresses this gap through better task instructions and demonstration selection. Our analysis shows that demonstration patterns change predicted record counts for unchanged inputs, suggesting that examples may constrain the model's interpretation of source content. Inspired by this finding, we propose SDA, a structured extraction framework that separates source-based semantic discovery from demonstration-guided annotation. Specifically, the framework comprises three modules. Semantic Discovery constructs natural-language candidates through complementary source readings without labeled demonstrations. Candidate Projection uses relevant precedents from an annotation memory to normalize candidates and delimit admissible annotation choices. Joint Compilation jointly resolves labels and record selection while preserving candidate meanings and previously fixed fields. Extensive experiments on three structured extraction tasks across five datasets show that SDA with GPT-4o-mini outperforms the strongest evaluated same-backbone LLM baseline on each dataset, with an average gain of 20.78 F1 points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.