acceptodds
Under review as a conference paper at ICLR 2027

SCTL: State Conditioned Temporal Ledgers for Long Horizon LLM Teams

Abstract

Long-horizon tasks give teams of large language model (LLM) agents many chances to drift from a cooperative plan. Existing methods fix each agent's share of the team's result before execution, although the work that remains changes as the task proceeds. We show that whether LLM agents stay on the plan depends on when each agent's share of the team's cost is charged, even when the total stays the same. We introduce the State-Conditioned Temporal Ledger (SCTL). At each checkpoint it estimates the cost that each group of agents would still bear from the observed state and splits the team's cost with the Shapley value. Each agent then sees how much it has paid and how much it still owes. Unlike classical time-consistent payment schedules, which are defined only along the agreed plan, SCTL makes each payment depend on the time and on the state. An agent that deviates keeps its unpaid balance as a penalty and pays it to the team when it returns. SCTL works on top of existing workflows without retraining. On nine long-horizon multi-agent LLM benchmarks and CoopEval, SCTL beats the strongest baseline fixed in advance in eight of ten settings, with a median long-horizon gain of 3.6%. On a controlled benchmark with exact values, it reduces the normalized gap in team cost by a median 37.1% relative to charging the same total in equal parts, and the advantage grows with the length of the task. Keeping the unpaid balance as a penalty cuts the cost gap of a delayed schedule from 0.611 to 0.166.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.