acceptodds
Under review as a conference paper at ICLR 2027

AhaBench: Do Agents Turn Experience into Reusable Insights? A Long-Horizon Benchmark for Continual Learning

Abstract

A solved example can help an agent complete a task without preparing it to solve the next one. AhaBench tests how agents use experience in later tasks and decisions. Aha-Puzzle tests questioning on a new mystery after solved episodes; Aha-Euler tests computation on new inputs after teaching; Aha-Vending follows decisions and feedback over a simulated business year. The benchmark leaves the learning mechanism open, and our evaluation is on fixed-weight agents. Earlier Aha-Puzzle solutions help some models, but scores remain far below solving with hints. Qwen 3.6 Plus gains +12.57 points over an attempt without earlier episodes. On single-step Aha-Euler tasks, it reaches 97.3% held-out accuracy with an explanation and code, versus 8.1% with the question and answer alone. In nominal Aha-Vending operation, Doubao Seed 2.0 Pro finishes at + $495, but its profit averaged over the year is -$10 because recovery arrives late. We release tasks, validators, a simulator, and evaluation records for studying how agents reuse experience on new tasks and sustain productive behavior over time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.