RATIO: A Reasoning Agent for Task-Adaptive Interaction Orchestration in Zero-Shot Object Navigation
Abstract
Zero-shot object navigation requires sequential decisions under partial observations. Existing methods increasingly use large language models (LLMs) for high-level reasoning, yet navigation capabilities are still commonly organized by predefined routines, while precise spatial computation is often mixed with semantic decision making. We introduce **RATIO** , a **R**easoning **A**gent for **T**ask-Adaptive **I**nteraction **O**rchestration for zero-shot object navigation. RATIO organizes geometric, semantic, and navigational information into an agent-facing environment representation and an executable navigation graph that the agent can query on demand, and exposes specialized navigation capabilities as tools. At each decision step, the agent selects tools based on the current task context, while dedicated modules perform geometric, retrieval, and candidate computations. By separating capability selection from execution, RATIO enables context-dependent navigation while preserving precise and reproducible spatial computation. We evaluate RATIO on HM3Dv1, HM3Dv2, MP3D, and HM3D-OVON Unseen, achieving a **53.0%** success rate and **34.0%** success weighted by path length on the open-vocabulary benchmark with GPT-5.6 Luna.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.