VisOrch: Benchmarking Agent-Designed Vision Pipelines
Abstract
Composing off-the-shelf vision models into task-specific pipelines avoids training from scratch, but manually designing them is costly. While typical vision agents invoke an LLM or VLM at every inference step, we study a deployment-efficient alternative: an agent designs an executable pipeline once during adaptation, which is then frozen and evaluated on unseen inputs with no such model calls at test time. We introduce VisOrch, a benchmark evaluating whether agents can automatically synthesize reusable vision pipelines from few examples. To ensure scores reflect genuine orchestration rather than a single dominant tool, VisOrch admits a task only when its strongest off-the-shelf component leaves substantial headroom. This yields nine tasks across four task families, totalling 540 adaptation, 480 validation, and 2,415 blind test examples; each task provides a fixed budget of 60 labelled images used to tune the tool baselines. On eight of the nine tasks, an agent-designed pipeline outperforms the strongest single tool on both metrics, whereas directly prompted frontier VLMs fall below it on every task they attempt. However, no single designer is universally reliable: iterative program search is ahead on both metrics at on four tasks, a general coding agent leads on three, both improve on DECIMER, and neither improves on both metrics for PlantSeg. Finally, building VisOrch exposed key evaluation pitfalls, including shortcut exploitation and degenerate metrics, which the benchmark protocol now explicitly guards against.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.