acceptodds
Under review as a conference paper at ICLR 2027

Risk-Controlled Speculative Prefetching for Tool-Augmented Dialogue

Abstract

Speculative decoding accelerates LLM inference by predicting future tokens. We apply the same principle to future user actions in tool-augmented dialogue: PRESTO (Predictive Speculative Tool Orchestration) predicts diverse hypotheses for the next user query, pre-executes read-only tools while the user reads the current response, and serves cached results on exact-key match when the actual query arrives, or discards with no degradation. We make three contributions. (1) The proxy reversal and its constructive converse. Offline embedding commit rate is a gameable proxy: predictors that maximize it degrade in-loop. The fix is coverage-WTA, adaptation supervised by the in-loop objective (labeling winners by realized cache hits); it beats the zero-shot predictor at equal planner budget, and no-training retrieval even when that retrieval is given a larger prefetch budget, in both -bench domains, reaching 70.1%/32.2% turn-level cache hits vs. 59.4%/23.7% zero-shot. (2) The entity-revelation boundary: an a-priori, per-domain ceiling on speculation (88.4% retail, 47.3% airline) that explains most remaining misses and scopes the magnitude, not the applicability, of adaptation's win. (3) A risk-controlled speculation gate: conformal risk control over a learned score operationalizing that boundary lifts fired precision to 82-87% (retail) and 50-59% (airline) over 69-70%/29-33% always-fire at the certified operating point; served-result correctness comes from exact-key matching, not the gate. A  600K-parameter MCL predictor (<1ms) enables the study: offline-stronger frontier predictors (82.6% vs. 54.4% commit) serve 0% in-loop at short prediction-time slack, where MCL serves 32-35%; with a 15s speculation window, MCL hides 32.6% of retail per-turn tool latency. Speculative prefetching must be evaluated, adapted, and gated in-loop.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.