Funnel: Public-Check-Guided Budget Allocation for Native Agent Harnesses
Abstract
Batch-oriented agent systems often improve task success by repeating complete trajectories, but uniform candidate budgets spend additional computation on both tasks that already pass a check and tasks that still require rescue. We introduce FUNNEL, a training-free, task-level budget scheduler that operates outside a native agent harness. FUNNEL runs one complete native trajectory for every task, then uses only protocol-permitted public-check outcomes to launch seven additional state-isolated trajectories for first-check failures. Candidates share neither mutable execution state nor feedback, and a fixed selector freezes the final candidate before independent verification. Across four configurations spanning two models, two harnesses, and two task families, Funnel-8 differs from fixed eight-candidate search by only -0.52 to +1.04 percentage points in final pass rate while reducing the number of generated candidates by 28.3–58.6% and reported GPU-hours by 18.2–48.2%. In matched Qwen3-32B/Hermes comparisons on Reasoning Gym and LiveCodeBench-v6, Funnel remains within -0.26 to +1.30 points of ASC and DSC while using fewer candidates and reported GPU-hours. A fixed-average-four UAB adaptation uses a slightly lower budget but trails Funnel by 3.12–7.03 points. These results show that concentrating attempts on first-check failures provides a competitive accuracy–cost operating point without changing native trajectories, while candidate savings do not imply proportional savings in GPU-hours.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.