acceptodds
Under review as a conference paper at ICLR 2027

Learning What to Perceive: Verifying and Consolidating Evidence Acquisition in Omni-Modal Agents

Abstract

Omni-modal active perception shifts multimodal understanding from static prediction to multi-turn acquisition of task-relevant evidence. Each perception action changes subsequent states and memory, making interaction part of task solving. Yet reinforcement learning still relies on final outcomes: successful trajectories may contain redundant or misleading acquisitions, while failed ones may already contain decisive evidence. Outcome-only supervision therefore obscures which turn, modality, and temporal interval actually helped. We introduce Modality-Attributed Turn-Level Credit Assignment (MATCA), a hierarchical credit framework that uses explicit media counterfactuals to assess whether the acquired media provides task-relevant evidence and whether the selected temporal interval is more informative than feasible alternatives. MATCA converts these comparisons into signed content and temporal-specificity credits and routes them only to responsible action-value fields, decoupling acquisition quality from trajectory success: useful actions in failed trajectories can be reinforced, while harmful ones in successful trajectories are suppressed. To overcome current-policy coverage limits, we introduce Selective Intervention Consolidation (SIC). SIC distills teacher actions into a reset student only when their student-completed branches outperform the corresponding student-action branches, thereby enriching the initialization for renewed MATCA optimization with beneficial acquisition actions. Under matched 7B settings, SIC-initialized MATCA outperforms OmniAgent by points on average across seven benchmarks. Despite using a smaller 7B backbone, it also surpasses the newer Qwen3-Omni-30B-A3B on four benchmarks: OmniVideoBench, LVOmniBench, WorldSense, and Video-MME.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.