Efficient Entity Disambiguation with Adaptive Candidate Pruning
Abstract
Entity disambiguation (ED) systems often use a two-stage pipeline: a candidate generator produces a fixed-size candidate set, and a more expensive candidate reranker selects the final prediction. However, because mention ambiguity varies substantially, reranking the same number of candidates for every mention wastes computation on easy cases. To address this issue, we propose a lightweight adaptive candidate pruning method that requires no additional model training. Inspired by conformal prediction, candidates are retained when their generator-score difference from the top candidate is below a threshold, yielding smaller sets for confident mentions and larger sets for ambiguous ones. We tune the threshold on development data for end-to-end accuracy, prioritizing practical utility over a formal coverage guarantee. Experiments on a standard ED benchmark show that our method reduces the average number of candidates passed to the reranker by 94% and achieves a 4.56 speedup, while improving accuracy. Further analyses demonstrate robustness to hyperparameter selection and show that the method places pruning boundaries close to the candidate depth required for each mention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.