acceptodds
Under review as a conference paper at ICLR 2027

Incentivizing Reasoning Effort in LLM Agents Without Ground-Truth Feedback

Abstract

Governing large AI swarms requires rewarding useful computation without verifying every answer. Multi-task peer prediction offers a route by paying agents from relationships among answers across tasks. To our knowledge, we provide the first empirical study of controllers choosing provider-supplied LLM reasoning settings under these payments. We model a game where utility is reward times peer score minus token cost and evaluate six models. A five-run online comparison finds no clear accuracy advantage over raw-agreement or flat payments. In a symmetric signal model, large rewards sustain maximum-effort equilibria in which agents systematically merge distinct answers.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.