ClaimCrawl: Citation-Supervised Retrieval and Reranking for Patent Prior-Art Search
Abstract
Patent prior-art search requires matching the technical content of a claim to earlier patents expressed in different terminology and lengthy descriptions. Beyond finding relevant documents, a useful search system must place them within the small set a searcher can inspect. Our analysis shows that high recall at deep retrieval cutoffs can coexist with limited target coverage among the earliest results. This gap motivates task adaptation at both retrieval and ranking stages, with separate evaluation of their contributions. We present ClaimCrawl, a framework that uses claim-level examiner citations to supervise a dense retriever and a pointwise language-model reranker. The retriever learns from cited positives and complementary lexical hard negatives and cross-device negatives. The reranker uses claim-conditioned excerpts to predict structured citation-alignment labels and scores, prioritizing targets within a fixed candidate pool. We also construct PAR4PC-Retrieve, linking PANORAMA queries and citation targets to a frozen corpus of over seven million U.S. patents. Supervised retrieval improves Recall@5000 by 8.71 percentage points over the same unadapted encoder. Within fixed candidate pools, reranking improves Recall@20 by 14.40 points over the retriever's original ordering and 8.56 points over task-adapted Qwen3-Reranker-8B. Model comparisons and ablations examine the contributions of task adaptation, and a blinded human study assesses the usefulness of the returned shortlists. Together, these results support citation-supervised retrieval and reranking as complementary components of patent prior-art search. Code: https://anonymous.4open.science/r/ClaimCrawl_code-BFBF/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.