ACE-Cap: Active Evidence Acquisition via Agentic Co-Evolution for Long-Paragraph Fine-Grained Audio Captioning
Abstract
Long-paragraph fine-grained audio captioning requires models to recover diverse acoustic facts while avoiding omissions and unsupported details. However, prevailing captioners remain passive one-shot generators: once a relevant detail is overlooked, they cannot identify the evidence gap, query the audio for targeted information, or adaptively determine when sufficient evidence has been collected. We formulate this task as active evidence acquisition and introduce Agentic Co-Evolution for Captioning (ACE-Cap), a framework that couples a Composer and an Instruct model through multi-turn interaction. A Captioner first produces an initial description. Conditioned only on this description and the interaction history, a text-only Composer asks targeted questions about unresolved acoustic attributes, while an audio-conditioned Instruct model provides grounded answers. The Composer then decides when to terminate the interaction and synthesizes the accumulated evidence into a final caption. ACE-Cap trains these roles through a unified gold-to-prediction reward derived from fixed, gold-grounded multiple-choice questions and a frozen caption-only judge. To assign credit within variable-length interactions, we introduce Leave-One-Out Per-turn Group Relative Policy Optimization (LOOP-GRPO). It replaces the trajectory-wide scalar advantage with span-aligned signals for individual questions, stopping decisions, and final synthesis. These signals measure leave-one-out evidence contributions, evidence sufficiency, and evidence preservation, respectively. Role-wise warm-up followed by alternating Composer and Instruct optimization keeps each update a well-defined single-policy problem while allowing the two roles to co-evolve. Extensive experiments across diverse audio benchmarks demonstrate state-of-the-art performance among evaluated models with open-weight backbones under a common caption-as-evidence protocol.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.