acceptodds
Under review as a conference paper at ICLR 2027

Canopy: Exploiting Piecewise-Smooth Tree Priors for Multi-Fidelity Bandits

Abstract

Many LLM inference problems - including model routing, prefix-cache management, prompt trimming, and test-time search - can be viewed as optimization over a tree. This structure arises naturally from autoregressive generation: every prefix define a node, and its continuations form a subtree below it. Internal nodes of the tree provide cheap but biased estimates of a region’s value, while leaf evaluations are expensive but accurate. Hierarchical bandit methods can exploit this structure, but typically require a specific smoothness schedule to be specified in advance, even though real objectives are often only piecewise smooth and their optima may lie near sharp boundaries. We introduce CANOPY, a multi-fidelity tree bandit that learns where the smoothness prior is valid rather than assuming it globally. CANOPY uses cheap random-path probes to construct an online certificate of local aggregation bias, then directs expensive leaf evaluations toward cells where the certificate detects a smoothness violation. We prove fixed-budget and regret guarantees whose additional cost is additive in the number of discontinuities, recovering the smooth-tree rate when no violations are present and approaching structure-blind search as violations become dense. Across routing, top-(k) identification, test-time search, caching, and prompt trimming, CANOPY consistently improves matched-budget performance, including higher top- recall on a -model pool, more SWE-bench Verified issues resolved than best-of-(N), and lower median time-to-first-token with prefix caching.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.