acceptodds
Under review as a conference paper at ICLR 2027

ArbiGraph: Composing Arbitrary Verifiable Tasks for Agent Evaluation

Abstract

Language model agents must often solve a sequence of problems and retain earlier answers as the workflow grows. Evaluating problems one at a time does not show how well agents retain and use earlier answers. To evaluate this ability, we intro- duce ARBIGRAPH, a framework for generating benchmark instances that combine problems with verifiable answers into graphs of dependent tasks. Each task is a problem, and each connection specifies an earlier answer that the agent must report to obtain the next problem’s input. An incorrect report returns an alternative input whose verified answer differs from the intended answer. This mechanism sup- ports tasks from different domains without requiring conversions between their input and output types. We test the framework with math problems and tasks that ask agents to predict the output of Python code. We evaluate three models on the same problems both in isolation and in three topologies, each containing 18 tasks. We construct an independent-error baseline by multiplying the isolated accuracies of the final task and all tasks it depends on. Averaged across the three topologies and the math and Python categories, measured accuracy falls below this baseline by 45.0 percentage points for Gemma-4-31B-it, 35.6 for Qwen3.8-27B, and 11.0 for GPT-OSS-120B. These additional losses indicate that solving problems within a larger workflow adds difficulty beyond the accumulation of errors on isolated problems. We also study how graph structure affects accuracy on the final task while keeping the number of tasks fixed.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.