acceptodds
Under review as a conference paper at ICLR 2027

CAIA: A Benchmark for Tool-Grounded Language Agents in Real-life Adversarial Cryptocurrency Tasks

Abstract

An agent answering a cryptocurrency question must find the relevant historical records and compute the requested quantity. Counting transactions, for example, does not answer a question about distinct wallets. We introduce CAIA, 178 expert-curated tasks with reference answers and checked tool sequences for obtaining them. We compare 17 models through a shared 23-tool interface, allowing up to five tool-use rounds within each run. Tools raise GPT-5's single-attempt success from 28.1% to 70.2%. Repeating the task in five independent runs produces a correct GPT-5 answer on 77.0% of tasks, but majority voting selects a correct answer on only 67.4%. This exposes an answer-selection bottleneck: correct candidates are available but lost during aggregation. Across 23,463 calls, an individual attempt's call count also has near-zero association with correctness. The interaction records connect these outcomes to source choice, historical scope, and calculation. We release the questions, expert reference routes, and records to support evidence-based analytical agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.