TRIAL: Benchmarking Temporal Rule Inference through Active Learning
Abstract
Agents that work on long-horizon tasks often have to learn how an unfamiliar system works by experimenting, and decide when their evidence is enough to act on. We introduce TRIAL (Temporal Rule Inference through Active Learning), a benchmark in which an agent identifies a hidden temporal rule, a deterministic finite automaton (DFA), by building experiments in a procedurally generated game world. An exact verifier tracks which rules remain consistent with the agent’s evidence, so TRIAL scores whether an answer is correct and also whether the evidence supports it. We show that frontier agents guess before their evidence settles the rule. When the rule space is large, they exhaust their budget first, and all 88 wrong answers in our main evaluation came within the last 27 of 240 decisions while a median of 260 candidate rules were still possible. When the rule space is small, they commit early with budget to spare, and halving the budget makes these guesses more common. Exact-answer scoring hides the guesses that happen to be right, such as the 29% of Opus’s correct answers with code that were made while other rules remained. With code, Astra and Opus write programs that enumerate the candidate rules, yet no system identifies a rule on the hardest condition. Without code, the agents remove more uncertainty than a greedy information-gain controller, but the uncertainty they report drifts away from what the evidence supports.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.