ProbeCompress: Compressing Long-Horizon Agent Benchmarks with Task-Derived Probes
Abstract
Evaluating AI agents on long-horizon benchmarks consumes many tokens and substantial execution time because each task can require multiple model calls, tool invocations, and environment interactions. Task-subset methods evaluate fewer tasks, but each retained task still runs to completion. We introduce ProbeCompress, which predicts a new agent's full-benchmark score from short probes derived from the benchmark's own tasks. Each probe is built from the public context of one task and checks one requirement with its own verifier. Using responses from calibration models, ProbeCompress selects a small probe suite and fits a score predictor; a new model then runs only the selected probes. Across six agent benchmarks with up to 18 models, ProbeCompress predicts full-benchmark scores at a prespecified operating point with a mean absolute error of 0.028, while using 222.1× fewer tokens and 45.7× less summed execution work than full evaluation. The strongest complete-task baseline in our comparison reaches similar error while using more than ten times as many tokens. An interface fitted on 14 earlier models and reused unchanged on 4 later models reaches a mean absolute error of 0.044 without refitting. These savings exclude the one-time cost of building and calibrating the probe suite and apply to the evaluation setup used for calibration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.