acceptodds
Under review as a conference paper at ICLR 2027

Visual Reasoning at the Harness Level: A Ladder for Controlled Evaluation

Abstract

Advances in foundation models are shifting coding and working agents from predefined tools toward general computational environments. Yet existing visual agents often couple model, harness, and training choices, making it difficult to determine how this transition affects visual reasoning. In this work, we introduce a unified evaluation framework to study how harness design allocates responsibility between the model and its environment. We instantiate this framework as a seven-rung harness ladder organized around expanding computational access and delegating state and evidence management. The ladder connects direct answering, predefined visual tools, generated code, and general workspaces under fixed models and tasks. We further construct (Visual Harness with Stratified Tasks for Agentic Reasoning), a benchmark of 600 tasks across five visual domains and four calibrated difficulty levels. Experiments across nine models show that code execution with a persistent kernel improves all four full-ladder models over fixed tools, whereas further workspace delegation offers no consistent advantage. Models with similar direct-answer performance benefit differently from computational access, motivating harness designs that expand computational freedom while retaining appropriate support for state and evidence. abstract

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.