acceptodds
Under review as a conference paper at ICLR 2027

Towards Human-AI Complementarity for AI Safety

Abstract

Safely overseeing AI is increasingly challenging for the humans doing it, as they must evaluate longer and more complex outputs with every new model generation. To mitigate this challenge, we could leverage complementarity: where humans and models make different mistakes, combining their judgments could yield an oversight signal stronger than either produces alone. We study this opportunity in realistic, safety-relevant settings, contributing the largest known suite of such tasks (3,957 instances across 8 datasets and 6 domains) as well as a platform and data pipeline for measuring human, AI, and human-AI performance under matched conditions. We collected 35,613 judgements from 416 human raters, finding a considerable opportunity for a complementarity effect, but also that the assistance methods we tested appeared to induce over-reliance rather than a synergy from coordinating human-AI teams. Confidence-based routing yields little improvement on average over the better single source. These results show that, while complementarity is an important and feasible engineering target for safe oversight, achieving it remains an open challenge for future work.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.