acceptodds
Under review as a conference paper at ICLR 2027

Beyond Query Similarity: Safety-Conditioned Cache Routing for LLM Response Reuse

Abstract

Semantic response caching reuses a cached response for a semantically similar query to improve LLM inference efficiency. However, query similarity alone can lead to unsafe reuse when safety conditions change. Existing approaches improve cache safety through stricter similarity checks, context-aware validation, or response verification, yet they may still miss safety conflicts between the cached response and the new request or introduce unnecessary verification overhead. We argue that safe and efficient semantic caching requires a reuse policy that explicitly represents the conditions required by a cached response and checks whether they remain satisfied by the new request. In this paper, we propose SafeCache, a safety-aware framework that separates candidate retrieval from safety approval and formulates semantic response reuse as a safety-conditioned routing problem. SafeCache represents cached-response reuse requirements and new-request safety facts as two safety condition sets and aggregates condition-level matches into an overall compatibility score. To prevent decisive conflicts from being diluted by aggregate matching, SafeCache checks explicit contradictions separately. SafeCache further estimates matching uncertainty and combines these safety signals with query similarity} to make the final reuse decision under safety and verifier-budget constraints. Experiments show that SafeCache achieves a better safety-utility trade-off among the evaluated baselines, enabling safer response reuse at comparable cache utility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.