acceptodds
Under review as a conference paper at ICLR 2027

IROHA: Long-Horizon Visual Reasoning with an Agentic Harness

Abstract

Visual reasoning asks a generative model to solve a task by generating its execution. Direct generation produces a candidate solution, but leaves open whether it satisfies the task and how to proceed if it fails. In long-horizon tasks, deciding what to do next also requires checking whether earlier actions have established the conditions needed by later ones. We introduce vIsual Reasoning Oriented Harness for Agents (IROHA) that makes this control explicit around frozen video and world-action models. Planning defines task objectives, while executable tools produce candidate visual solutions and first-frame guidance for structured tasks, or visual grounding for natural scenes. Verification checks whether generated outcomes realize these requirements; its feedback guides candidate selection or revision of the task analysis and generation setup. With frozen state-of-the-art video and world-action backbones, IROHA improves the VBVR score from 0.65 to 0.78 and LIBERO-Plus simulator success rate from 72.8% to 74.8%. It also improves the RBench scores across three backbones, including gains on long-horizon instructions. Intriguingly, our RBench analysis shows substantial gains in task accomplishment alongside nearly unchanged rendering consistency. We further demonstrate how IROHA assembles multistep video solutions to the challenges like Hanoi Tower in real-world style scenes, using state checks, retries, and backtracking. These findings show that explicit reasoning and feedback unlock the potential of existing generative capabilities in complex tasks, without additional training or parameter updates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.