acceptodds
Under review as a conference paper at ICLR 2027

DCAgent: Towards Dense Visual Perception via a Divide-and-Conquer Tool-Use Agent

Abstract

Frontier Vision-Language Models (VLMs) collapse on dense visual perception—tasks that require enumerating or localising large numbers of small, tightly packed objects. The structural cause is a fixed-budget visual encoder that smears small dense objects before they reach the language model. The intuitive remedy—divide-and-conquer crops—has been explored as a hard-coded external pipeline, but never as a learned, agentic capability inside the VLM itself. We present DCAgent, a tool-use agent that natively learns to divide and conquer its own image input through two visual tools, Divide and Merge: Divide splits the current view into sub-pieces the agent itself selects, and Merge aggregates the resulting leaf detections via class-aware non-maximum suppression. The multi-turn tool-use policy is trained with GRPO under an end-to-end final reward combining count MAE and a DETR-style Hungarian-matched detection IoU on the merged output. Three task-specialised agents are trained on SKU-110K, SOD-Aerial, and LVLM-Count from a Qwen3-VL-8B starting point and each surpasses frontier proprietary VLMs on its target dataset. We then consolidate them into a single deployable model with Online Policy Distillation under a four-teacher setup. The merged model has internalised the use of the divide-and-merge tools, not the patterns of any single dataset: it auto-triggers the tools on counting and dense-perception inputs and bypasses them on knowledge-and-reasoning inputs, so the single checkpoint substantially improves on all three training datasets, generalises to CountBench, and maintains the Qwen3-VL-8B base on MMMU, MMMU-Pro, MMStar, and MMBench-V1.1.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.