acceptodds
Under review as a conference paper at ICLR 2027

DIVE: Making Judgment a First-Class Capability in Long-Horizon Agents

Abstract

Agents working on long tasks must decide when to continue, check their work, change direction, or stop. A final task score gives little information about the value of these decisions. We introduce DIVE, which encodes judgment procedures as reusable Cognitive Judgment Skills. Each Skill specifies when prior experience applies, what evidence a decision requires, and when to reconsider it. A Detect–Invoke–Verify loop connects these procedures to the host agent without replacing its action generator. We evaluate DIVE in software repair and scientific program search, using task outcomes and continuations from the same state. Guarded stopping reduces aggregate Target tokens by 72.52% across 16 matched search pairs. Every pair passes its task-specific final-quality check, though five have lower anytime quality. In the MiniMax and DeepSeek repair summaries, Jev uses fewer reported tokens than Native or DIVE with Target-based judgment. On 30 SWE-bench Pro tasks, outer-loop DIVE with Codex SDK and GPT-5.6-Luna resolves 27 tasks versus 24 for Native, using 19.41% fewer input-plus-output tokens. The decision-level results are less consistent. Reuse has positive mean local effects in four cohorts, but selection gains and five-step outcomes vary. DIVE provides a way to measure these decisions and their costs; the observed task-level benefits do not establish a general advantage for its selection policy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.