acceptodds
Under review as a conference paper at ICLR 2027

When Good Teammates Are Weaker Alone: Handoff Regret in Skill-Compatible AI

Abstract

A strong agent that shares control with a weaker partner can leave behind states the partner cannot handle, even when each of its own actions is sound. We study this through handoff regret, the value lost when a fixed partner rather than the agent takes the next action. The central object is the handed-over loss: the solo losses of the partner's actions, summed over its observation classes with the team's weights. Team loss relative to the solo optimum is exactly the agent's own error plus this loss averaged under the partner's rule; the least any rule on the same observations can pay is the value of the information the partner lacks, and the rest is its decision error. Refinement, as in Blackwell's comparison of experiments, is exactly the order under which the information term never rises in any task, but behaviour identifies the split only up to an interval whose ends respond to different interventions, and a construction at one end shows that the team gain per unit of solo loss can be unbounded. The same accounting yields partner-aware corrections to an existing value oracle with explicit error terms. In tag-team chess with Maia-2 partners and strict alternation, the correction beats an engine given the same search budget head to head (57.2% of points), while the extra search alone gains nothing. In navigation, interventions on the agent's planning, the partner's rule and its information each move the term the decomposition predicts; on human games, estimated handoff regret predicts losing moves better than standard difficulty measures, and adds to them.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.