acceptodds
Under review as a conference paper at ICLR 2027

CARD: Calibrated Adaptive Reliance Decoding in LLM-Augmented Search

Abstract

Search engines augmented with large language models (LLMs) synthesize retrieved web content into summarized responses. When handling open-ended queries, LLMs implicitly select and rank candidate items, effectively functioning as recommenders. However, the LLM's reliance on parametric versus retrieved knowledge during this process is difficult to control: over-reliance on parametric knowledge leads to systematic favoritism toward known content, while over-reliance on retrieved content makes it vulnerable to manipulation of the retrieved documents. To investigate this problem, we first define (PRS) to quantify the strength of parametric knowledge reliance in LLM recommendation rankings. We validate this metric through a Latin-square-based controlled experiment across five LLMs, confirming systematic parametric preference. We then propose CARD (alibrated daptive eliance ecoding), a pluggable inference-time method that leverages PRS to enable bidirectional adaptive control over knowledge reliance. CARD measures KL divergence of token probability distributions after context insertion to derive an adaptive intervention strength on external knowledge, and calibrates the anchoring weight of the base predictions using the model-level PRS. Experiments on three open-source LLMs show that CARD mitigates parametric knowledge bias, narrowing the ranking gap between parametric and non-parametric items by 55–60%, and defends against adversarial attacks on retrieved content—suppressing retrieval poisoning to near-random rankings and neutralizing up to 95% of black-hat search engine optimization (SEO) attack effects. Our code is available at https://anonymous.4open.science/r/CARD-1A64.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.