NAVIGATE: Next-Action Verifiable and Flexible Incentives with Guided Adaptive Tool Execution for Multimodal Reasoning
Abstract
Multimodal Large Language Models (MLLMs) remain brittle on tasks requiring multi‑step perception–reasoning and tool use. Existing approaches degrade perceptual accuracy, overfit to tool‑call patterns, or generalize poorly to multi‑step rollouts—manifesting as irrelevant or skipped tool invocations. We present NAVIGATE, a multi‑stage framework that teaches MLLMs to plan next actions, call perception specialists, and verify outcomes end‑to‑end. The first stage performs tool‑aware adaptive sampling: a curriculum that weighs trajectories by tool usage, compositionality, and perceptual hardness, progressing from easy grounding to complex multi‑tool reasoning. The second stage augments supervised fine‑tuning (SFT) with a planner head that forecasts action sequences and regresses to action embeddings to respect similarity among tools. The final stage applies reinforcement learning (RL) with verifiable (terminal accuracy, format) and flexible (step‑wise consistency, embedding‑transport discrepancy, variance) rewards to optimize interactive rollouts. On multiple benchmarks, NAVIGATE yields around 10% gains over SFT and RL baselines across different backbones. We also conduct a series of ablation experiments and analyses to gain deeper insight into our framework.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.