acceptodds
Under review as a conference paper at ICLR 2027

Towards Multimodal Agents with Cognitive Modality Planning and Reflective Reasoning

Abstract

Multimodal agents must decide which perceptual modality to read, when to stop refining their belief, and how to keep an answer honest about its own uncertainty. Existing pipelines either run a single fixed toolchain (no planning), or invoke every modality and then reflect (no budget), or rely on a heuristic LLM-as-judge for confidence (no calibration). We intro duce CoMPR (Cognitive Modality Planning + Reflective Reasoning), a training-friendly framework that formulates modality selection as a bud geted submodular maximization, runs a short chain of reflective revise and-verify passes with a temperature-calibrated verifier, and is post-trained with GRPO under a composite reward that explicitly rewards evidence grounded, calibrated, budget-efficient behaviour. CoMPR + a GRPO warm-start (CoMPR+RL) is evaluated on five multimodal QA benchmarks and two agentic benchmarks across four backbone LLMs, and compared against eleven strong baselines including CoT, self-consistency, ReAct, Re flexion, Self-Refine, Critic, Set-of-Mark, random planning, all-modalities, and the recent CogPlanner and DeepEyes agents. On Qwen2.5-VL-7B CoMPR+RL improves mean accuracy by +5 .0 over the strongest base line, reduces expected calibration error from 0.148 to 0.031, and shortens latency by ∼17%, while remaining backbone-agnostic and competitive with the much larger GPT-4o.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.