Under review as a conference paper at ICLR 2027
Incentivizing Reasoning Effort in LLM Agents Without Ground-Truth Feedback
Abstract
Governing large AI swarms requires rewarding useful computation without verifying every answer. Multi-task peer prediction offers a route by paying agents from relationships among answers across tasks. To our knowledge, we provide the first empirical study of controllers choosing provider-supplied LLM reasoning settings under these payments. We model a game where utility is reward times peer score minus token cost and evaluate six models. A five-run online comparison finds no clear accuracy advantage over raw-agreement or flat payments. In a symmetric signal model, large rewards sustain maximum-effort equilibria in which agents systematically merge distinct answers.
open until 14 Dec 2026
est. 32% chance this paper gets accepted at ICLR 2027.
Reject 68%Accept 32%
What do you think this paper will get?
All positions stay anonymous.
Related papers
Loading the map…
Discussion (0)
Sign in to comment.